Decoding Public Sentiment: What GDELT and Twitter Reveal About Online Education During COVID-19
Opinion Mining about Online Education Basing on GDELT and Twitter Data
This study presents a multi-platform opinion mining framework using GDELT and Twitter data to analyze public sentiment and media discourse regarding online education during the COVID-19 pandemic. By integrating GDELT Summary tools with Python-based LDA modeling, the research reveals the misalignment between institutional media focus and individual concerns.
TL;DR
The COVID-19 pandemic forced a global pivot to online learning, creating a massive digital footprint of public opinion. This research leverages the GDELT Project and Twitter API to map the landscape of media narratives versus public reality. The study finds that while media volume drives search behavior, there remains a significant disconnect between what the news reports and what students actually feel.
Problem & Motivation: The PHEIC Knowledge Gap
Public Health Emergencies of International Concern (PHEIC) create volatile informational environments. Historically, understanding educational public opinion was restricted to surveys or small-scale qualitative studies. The authors argue that these methods are too slow for an active pandemic. Their goal was to use "Big Data" to identify the incubation, burst, and dissipation phases of public concern to help institutions adapt more effectively.
Methodology: A Dual-Track Mining Framework
The researchers combined two powerful data sources to capture both the "macro" (News Media) and "micro" (Social Media) perspectives.
1. Macro Analysis (GDELT)
GDELT (Global Database of Events, Language, and Tone) was used to monitor global news coverage.
- Volume & Tone Timelines: Tracking when online education peaked in the news and whether the narrative was positive or negative.
- Visual Narratives: Using Google’s Cloud Vision API to analyze image tags and topics associated with online education news.
2. Micro Analysis (Twitter & Python)
To capture the "pulse of the people," the authors scraped Twitter:
- LDA Topic Modeling: A generative statistical model used to discover the latent "hidden" themes in thousands of tweets.
- Word Frequency: Identifying the most common terminology used by netizens.
Fig. 1: Volume Timeline showing the flux in media attention toward online education.
Key Results: Media Stimuli vs. Netizen Concerns
The study’s most compelling finding is the thematic misalignment between different actors:
- The Media's Focus: News outlets centered on systemic issues, urban planning, and how societies can rearrange space for education.
- The Public's Focus: Netizens on Twitter were far more concerned with personal experience, the "virtual" nature of the classroom, and localized school issues.
- The Search Link: There is a clear relationship between media volume and Google search trends, confirming that media remains a primary driver for Health Information Seeking Behavior (HISB).
Table 1: LDA based topic modeling revealing latent concerns such as "virtual" environments and "focus".
Critical Insight & Future Outlook
The study highlights a critical "Communication Gap." Media propaganda often misses the mark of audience interest. For online education to succeed in the long term—as part of a lifelong education system—institutions must listen to the sentiment of the learners, not just the logistics of the providers.
Limitations
The study primarily relies on English-language Twitter data and GDELT’s broad labels. Future research could benefit from:
- Sentiment Granularity: Moving beyond "Positive/Negative" to nuanced emotions like anxiety or Zoom fatigue.
- Cross-Lingual Mining: Comparing how different cultures (e.g., China vs. US) reacted to the online transition.
Conclusion
This paper serves as a blueprint for using large-scale open-source datasets to monitor social changes in real-time. By bridging the gap between media narratives and public sentiment, educators can better design digital environments that resonate with the actual needs of the students.
