[Paper Review] FakeCovid -- A Multilingual Cross-domain Fact Check News Dataset for COVID-19
Introduces FakeCovid, a multilingual cross-domain dataset of 5,182 fact-checked COVID-19 articles from 92 sites in 40 languages across 105 countries, and provides a baseline classifier achieving an F1 of 0.76 for detecting false news.
In this paper, we present a first multilingual cross-domain dataset of 5182 fact-checked news articles for COVID-19, collected from 04/01/2020 to 15/05/2020. We have collected the fact-checked articles from 92 different fact-checking websites after obtaining references from Poynter and Snopes. We have manually annotated articles into 11 different categories of the fact-checked news according to their content. The dataset is in 40 languages from 105 countries. We have built a classifier to detect fake news and present results for the automatic fake news detection and its class. Our model achieves an F1 score of 0.76 to detect the false class and other fact check articles. The FakeCovid dataset is available at Github.
Motivation & Objective
- Create a large multilingual dataset of fact-checked COVID-19 news from multiple domains and sources.
- Annotate articles into content-based categories to enable cross-domain analysis.
- Provide a baseline model for automatic fake-news detection across languages.
- Enable analysis of cross-language and cross-country characteristics of COVID-19 misinformation.
Proposed method
- Collect 5,182 fact-checked COVID-19 articles from 92 fact-checking websites.
- Reference sources from Poynter and Snopes to curate articles (04/01/2020–15/05/2020).
- Manually annotate articles into 11 content categories.
- Ensure multilingual coverage across 40 languages and 105 countries.
- Train a classifier to detect fake news and report its performance on the false class and other fact-check classes.
Experimental results
Research questions
- RQ1How does a multilingual, cross-domain dataset of COVID-19 fact-checked news look across languages and countries?
- RQ2Can a classifier trained on this dataset effectively distinguish fake news from other fact-checked content across languages?
- RQ3What are the distribution and characteristics of the 11 content categories across languages and domains?
Key findings
- The dataset contains 5,182 fact-checked COVID-19 articles.
- Articles come from 92 fact-checking websites.
- The data covers 40 languages across 105 countries.
- A classifier trained on FakeCovid achieves an F1 score of 0.76 for detecting the false class and other fact-check articles.
- The dataset is available on GitHub for public use.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.