[Paper Review] Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset
The paper presents CMU-MisCOV19, a large annotated Twitter dataset and analyses two competing COVID-19 misinformation communities (informed vs misinformed) across network structure, sociolinguistics, and disinformation participation.
From conspiracy theories to fake cures and fake treatments, COVID-19 has become a hot-bed for the spread of misinformation online. It is more important than ever to identify methods to debunk and correct false information online. In this paper, we present a methodology and analyses to characterize the two competing COVID-19 misinformation communities online: (i) misinformed users or users who are actively posting misinformation, and (ii) informed users or users who are actively spreading true information, or calling out misinformation. The goals of this study are two-fold: (i) collecting a diverse set of annotated COVID-19 Twitter dataset that can be used by the research community to conduct meaningful analysis; and (ii) characterizing the two target communities in terms of their network structure, linguistic patterns, and their membership in other communities. Our analyses show that COVID-19 misinformed communities are denser, and more organized than informed communities, with a possibility of a high volume of the misinformation being part of disinformation campaigns. Our analyses also suggest that a large majority of misinformed users may be anti-vaxxers. Finally, our sociolinguistic analyses suggest that COVID-19 informed users tend to use more narratives than misinformed users.
Motivation & Objective
- Create a diverse, annotated COVID-19 Twitter dataset with a comprehensive codebook for misinformation analysis.
- Characterize informed and misinformed communities by network structure, linguistic patterns, and disinformation affiliation.
- Assess bot involvement and vaccination stance within misinformed groups.
- Provide data and methods to facilitate reproducible misinformation detection research.
Proposed method
- Collect Twitter data using diverse COVID-19 keywords and hashtags across three collection dates.
- Annotate tweets into 17 categories describing misinformation and true information categories.
- Compute user valence from tweet annotations to assign users to informed or misinformed groups.
- Augment data with user timelines filtered to COVID-19 related tweets for network, bot, and sociolinguistic analyses.
- Construct retweet, mention, and reply networks and compute network density for each group.
- Identify bots with Bot-Hunter and compare proportions between groups using a two-sample z-test.
- Perform LIWC-based sociolinguistic analysis on non-bot timelines to compare lexical categories between groups.
- Assess vaccination stance within misinformed users via hashtag valence propagation.
Experimental results
Research questions
- RQ1Can a diverse annotated Twitter dataset be created for COVID-19 misinformation analysis?
- RQ2What are the network structural differences between informed and misinformed COVID-19 communities?
- RQ3Do misinformed communities show higher bot involvement than informed communities?
- RQ4Are there distinct sociolinguistic patterns between the two groups, such as narrative usage or formality?
- RQ5What is the vaccination stance distribution within the misinformed group, and how do bots factor in?
Key findings
- Misinformed communities are denser than informed ones, indicating stronger echo chambers.
- About 47% of users are informed, 29% misinformed, and 24% ambiguous or irrelevant.
- Bots constitute 19% of misinformed users versus 11% of informed users, a statistically significant difference.
- Informed users exhibit more narrative language, higher pronoun and function word usage, and greater authenticity.
- Both groups show negative overall tone, with misinformed users tending toward higher informal language (inconclusive).
- Within misinformed users posting vaccine-related content, 41% are anti-vaxxers and 22% are pro-vaxxers (37% ambiguous).
- Disinformation campaigns are suggested by higher bot presence in misinformed groups.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.