[Paper Review] Dataset of Fake News Detection and Fact Verification: A Survey
This paper presents a comprehensive survey of 118 datasets across fake news detection, fact verification, and related tasks, systematically categorizing their characteristics, utilization tasks, and limitations. It identifies key challenges in dataset construction—such as data drift, bias, and lack of benchmarks—and proposes research opportunities to improve model robustness and reproducibility in fake news research.
The rapid increase in fake news, which causes significant damage to society, triggers many fake news related studies, including the development of fake news detection and fact verification techniques. The resources for these studies are mainly available as public datasets taken from Web data. We surveyed 118 datasets related to fake news research on a large scale from three perspectives: (1) fake news detection, (2) fact verification, and (3) other tasks; for example, the analysis of fake news and satire detection. We also describe in detail their utilization tasks and their characteristics. Finally, we highlight the challenges in the fake news dataset construction and some research opportunities that address these challenges. Our survey facilitates fake news research by helping researchers find suitable datasets without reinventing the wheel, and thereby, improves fake news studies in depth.
Motivation & Objective
- To provide a large-scale, systematic review of 118 datasets in fake news research, covering detection, fact verification, and related tasks.
- To address the gap in existing surveys by focusing specifically on dataset characteristics, tasks, and limitations rather than only on detection and verification methods.
- To highlight critical challenges in dataset construction—such as data drift, bias, and lack of stable benchmarks—that hinder reproducible and robust model evaluation.
- To guide researchers in selecting appropriate datasets without reinventing the wheel, thereby accelerating and deepening fake news research.
- To identify underexplored research opportunities in dataset design, including dynamic knowledge integration and bias mitigation.
Proposed method
- Conducted a large-scale survey of 118 datasets from three domains: fake news detection (51), fact verification (25), and other tasks (42), including satire and misinformation analysis.
- Categorized datasets based on source (e.g., social media, news outlets, fact-checking websites), annotation type (binary, multi-class, sequence labeling), and temporal scope.
- Analyzed dataset characteristics such as language, domain coverage, label reliability, and availability of metadata (e.g., user profiles, post timestamps).
- Evaluated datasets for benchmark suitability by assessing data stability, API dependency, and temporal consistency—especially critical for social media-based datasets.
- Identified and discussed biases in datasets, including word-level attention bias, annotator bias, political bias, and gender/racial bias, using models trained on FNC and FEVER datasets as case studies.
- Proposed solutions such as dynamic knowledge graph integration and Wikidata-based proper noun replacement to improve model robustness against data drift.
Experimental results
Research questions
- RQ1What are the key characteristics and differences among 118 datasets in fake news detection, fact verification, and related tasks?
- RQ2Why are existing social media-based datasets unsuitable as stable benchmarks for model evaluation?
- RQ3How do biases—such as word-level attention bias, annotator bias, and political bias—affect model performance and fairness in fake news detection?
- RQ4To what extent does data drift, particularly due to changing political or social contexts (e.g., new presidents, emerging events), degrade model generalization?
- RQ5What research opportunities exist to improve dataset quality and model robustness in fake news research?
Key findings
- The survey identifies 118 datasets across three main categories: 51 for fake news detection, 25 for fact verification, and 42 for other tasks such as satire and misinformation analysis.
- Many widely used datasets like Twitter15, Twitter16, and FakeNewsNet are not stable benchmarks due to reliance on social media APIs and dynamic data, leading to inconsistent results over time.
- Bias in datasets—especially word-level attention bias and annotator bias—significantly affects model generalization, with models in FNC and FEVER datasets showing over-reliance on specific noun phrases.
- Data drift due to temporal changes in topics and proper nouns (e.g., new political figures) reduces model performance in future or cross-domain settings, as seen in models trained on 2017 data failing on 2021 content.
- The NELA-GT dataset demonstrates a viable approach to long-term stability by updating content annually, suggesting a model for sustainable dataset maintenance.
- The study reveals a critical research gap: no widely accepted benchmark exists for fake news detection, and future work must prioritize stable, dynamic, and bias-aware dataset construction.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.