[Paper Review] COVID-19 Twitter Dataset with Latent Topics, Sentiments and Emotions Attributes
This paper presents a large-scale, globally representative Twitter dataset (252M tweets, Jan 2020–Jun 2022) with fine-grained, tweet-level annotations for latent topics, sentiment valence, and four core emotions (fear, anger, sadness, happiness). Using LDA for topic modeling and the CrystalFeel model for emotion and sentiment scoring, the dataset enables multidisciplinary research in public health, psychology, and social science by providing rich, semantically meaningful attributes for real-time analysis of pandemic-related discourse.
This paper describes a large global dataset on people's discourse and responses to the COVID-19 pandemic over the Twitter platform. From 28 January 2020 to 1 June 2022, we collected and processed over 252 million Twitter posts from more than 29 million unique users using four keywords: "corona", "wuhan", "nCov" and "covid". Leveraging probabilistic topic modelling and pre-trained machine learning-based emotion recognition algorithms, we labelled each tweet with seventeen attributes, including a) ten binary attributes indicating the tweet's relevance (1) or irrelevance (0) to the top ten detected topics, b) five quantitative emotion attributes indicating the degree of intensity of the valence or sentiment (from 0: extremely negative to 1: extremely positive) and the degree of intensity of fear, anger, sadness and happiness emotions (from 0: not at all to 1: extremely intense), and c) two categorical attributes indicating the sentiment (very negative, negative, neutral or mixed, positive, very positive) and the dominant emotion (fear, anger, sadness, happiness, no specific emotion) the tweet is mainly expressing. We discuss the technical validity and report the descriptive statistics of these attributes, their temporal distribution, and geographic representation. The paper concludes with a discussion of the dataset's usage in communication, psychology, public health, economics, and epidemiology.
Motivation & Objective
- To develop a comprehensive, globally representative dataset of public discourse on the COVID-19 pandemic from Twitter.
- To address the lack of publicly available, fine-grained, tweet-level annotations for topics, sentiments, and emotions in pandemic-related social media data.
- To enable multidisciplinary research by providing rich, semantically and psychologically meaningful attributes for each tweet.
- To support longitudinal analysis of public sentiment, emotional trends, and topic evolution across countries and over time.
Proposed method
- Collected over 252 million English-language tweets using Twitter's standard API with keywords: 'corona', 'wuhan', 'nCov', 'covid', and later 'vaccine' related terms.
- Applied Latent Dirichlet Allocation (LDA) to identify and label ten dominant latent topics per tweet.
- Used the pre-trained CrystalFeel model to predict sentiment valence (0–1) and intensity scores (0–1) for fear, anger, sadness, and happiness emotions.
- Generated categorical sentiment (very negative to very positive) and dominant emotion (fear, anger, sadness, happiness, no specific emotion) labels from the quantitative scores.
- Extended data coverage to 28 January 2020 – 1 June 2022, including 8 months of vaccine-related content from November 2021.
- Ensured compliance with Twitter’s Terms of Service and applied a CC BY-NC 2.0 license with restrictions for commercial use.
Experimental results
Research questions
- RQ1How do public sentiments and emotional intensities (fear, anger, sadness, happiness) evolve over time during the COVID-19 pandemic across different countries?
- RQ2What are the dominant latent topics in public discourse on Twitter, and how do they shift in response to public health events and policy changes?
- RQ3To what extent do sentiment and emotion patterns correlate with epidemiological trends or government interventions?
- RQ4How do emotional and sentiment profiles differ between users in different countries, and what cultural or contextual factors may explain these differences?
- RQ5Can the intensity and type of emotion expressed in tweets predict broader public mental health trends or behavioral responses?
Key findings
- The dataset contains over 252 million tweets collected from 29 million unique users between January 28, 2020, and June 1, 2022, with full coverage extended to 30 representative countries.
- The LDA model successfully identified ten dominant latent topics, with each tweet assigned a binary relevance score to each topic cluster.
- Sentiment valence scores ranged from 0 (extremely negative) to 1 (extremely positive), with average values showing a shift from predominantly negative to mixed or positive sentiment over time.
- Emotion intensity scores for fear, anger, sadness, and happiness were quantified on a 0–1 scale, enabling longitudinal analysis of emotional dynamics.
- The dataset includes 8 months of vaccine-related content from November 3, 2021, to June 1, 2022, allowing for analysis of public sentiment and emotional responses to vaccination rollouts.
- The dataset has been used in multiple studies, including one showing that future orientation is associated with higher joy and anger, while past orientation correlates with fear and sadness.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.