Skip to main content
QUICK REVIEW

[Paper Review] A Large-Scale Comparative Study of Accurate COVID-19 Information versus Misinformation

Yida Mu, Ye Jiang|arXiv (Cornell University)|Apr 10, 2023
Misinformation and Its Impacts4 citations
TL;DR

This study conducts a large-scale comparative analysis of 242 million COVID-19 tweets to identify distinct characteristics of accurate information versus misinformation, using a newly created, evidence-based misinformation classification dataset. The results show that misinformation spreads 158% faster than accurate content, is highly correlated with negative emotions and conspiracy themes, and is disproportionately removed by platforms—especially those related to conspiracy theories—while origin-related misinformation remains largely unaddressed.

ABSTRACT

The COVID-19 pandemic led to an infodemic where an overwhelming amount of COVID-19 related content was being disseminated at high velocity through social media. This made it challenging for citizens to differentiate between accurate and inaccurate information about COVID-19. This motivated us to carry out a comparative study of the characteristics of COVID-19 misinformation versus those of accurate COVID-19 information through a large-scale computational analysis of over 242 million tweets. The study makes comparisons alongside four key aspects: 1) the distribution of topics, 2) the live status of tweets, 3) language analysis and 4) the spreading power over time. An added contribution of this study is the creation of a COVID-19 misinformation classification dataset. Finally, we demonstrate that this new dataset helps improve misinformation classification by more than 9\% based on average F1 measure.

Motivation & Objective

  • To understand the statistical and linguistic differences between accurate and false COVID-19 information on social media.
  • To address the challenge of misinformation during the pandemic by identifying distinct patterns in topic distribution, language use, and spreading dynamics.
  • To develop and validate a new, high-quality dataset for classifying COVID-19 misinformation to improve detection models.

Proposed method

  • The study collected over 242 million COVID-19-related tweets via the Twitter API using a curated list of credible sources and keywords.
  • A novel, evidence-based misinformation classifier was trained using a newly created dataset enriched with manual annotations and post-processing to improve reliability.
  • Topic modeling and keyword-based filtering were used to identify and categorize misinformation into thematic clusters, including conspiracy theories and public authority criticism.
  • Linguistic analysis was conducted using Bag-of-Words and LIWC (Linguistic Inquiry and Word Count) to compare emotional tone, word frequency, and semantic categories between misinformation and accurate content.
  • Spreading power was measured by tracking retweet and engagement dynamics over time, including duration and peak velocity.
  • The performance of the new classifier was evaluated using leave-claim-out cross-validation, demonstrating a significant improvement in F1 score over prior state-of-the-art models.
A Large-Scale Comparative Study of Accurate COVID-19 Information versus Misinformation

Experimental results

Research questions

  • RQ1What are the differences in topics and languages between accurate information and misinformation?
  • RQ2What types of misinformation have social media platforms addressed, and how does this compare to the actual spread of misinformation?
  • RQ3What is the spreading power of different types of misinformation over time, and how does it compare to accurate information?

Key findings

  • Misinformation spreads 158% faster than accurate information, with conspiracy theory-related content showing the longest spreading duration, remaining active for over 32 hours after initial posting.
  • Conspiracy theory-related misinformation accounts for approximately one-third of all misinformation sources, and is the primary target of platform moderation, with over 40% of such tweets removed.
  • Tweets related to the origin of the virus received the least platform intervention, with nearly 70% still accessible after publication.
  • Linguistic analysis using LIWC revealed that misinformation is strongly associated with 'anger', 'negative emotion', and 'death' categories, while accurate information is linked to 'positive emotion', 'authenticity', and 'social' terms.
  • The Bag-of-Words analysis identified keywords such as '#nwoevilplans', '#chinaliedandpeopledied', and '#plandemic' as highly indicative of misinformation.
  • The newly created dataset improved misinformation classification performance by more than 9% in average F1 measure, achieving an F1 score of 0.70 compared to the baseline of 0.51.
Figure 1: (a) : Screenshot of an IFCN debunk post. The post includes 1) fact checking organisation, 2) misinformation claim, 3) explanation of why the claim is false and 4) the link to full debunk article. (b) : Partial screenshot of full debunk article of the IFCN debunk post.
Figure 1: (a) : Screenshot of an IFCN debunk post. The post includes 1) fact checking organisation, 2) misinformation claim, 3) explanation of why the claim is false and 4) the link to full debunk article. (b) : Partial screenshot of full debunk article of the IFCN debunk post.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.