[Paper Review] NELA-GT-2018: A Large Multi-Labelled News Dataset for The Study of Misinformation in News Articles
Introduces a large, engagement-independent dataset of 713,534 English-language news articles from 194 outlets (02/2018–11/2018) with multi-site ground-truth labels for veracity and credibility.
In this paper, we present a dataset of 713k articles collected between 02/2018-11/2018. These articles are collected directly from 194 news and media outlets including mainstream, hyper-partisan, and conspiracy sources. We incorporate ground truth ratings of the sources from 8 different assessment sites covering multiple dimensions of veracity, including reliability, bias, transparency, adherence to journalistic standards, and consumer trust. The NELA-GT-2018 dataset can be found at https://doi.org/10.7910/DVN/ULHLCB.
Motivation & Objective
- Fill the gap of large, multi-dimensional ground-truth labels for misinformation research independent of social media engagement.
- Provide a long-span dataset suitable for machine learning and qualitative studies of news veracity and bias.
- Corroborate source-level ground truth across multiple assessment platforms to enable diverse analyses.
Proposed method
- Systematically crawled RSS feeds of 194 news outlets twice daily from 02/2018 to 11/2018 to collect 713,534 articles.
- Compiled ground-truth labels for sources from eight assessment sites covering reliability, bias, transparency, and trust.
- Organized article data in an SQLite database with date, source, title, and cleaned text; labels stored in CSV per source.
- Compared and described the eight ground-truth sources (NewsGuard, Pew, Wikipedia, OpenSources, MBFC, AllSides, BuzzFeed News, PolitiFact) and their labeling schemes.
- Provided a documented dataset with accompanying use-case guidance for distant supervision and semi-supervised learning.
- Proposed that the dataset supports ground-truth-driven machine learning and mixed-method research into misinformation tactics.
Experimental results
Research questions
- RQ1How can a large, multi-label ground-truth dataset enable robust misinformation research independent of engagement signals?
- RQ2What are the characteristics and coverage of veracity and bias across a diverse set of news sources over an extended time period?
- RQ3How can source-level ground truth from multiple assessment sites be leveraged for article-level analyses and model training?
Key findings
- The dataset contains 713,534 articles from 194 news outlets collected between 02/2018 and 11/2018.
- Ground truth labels are corroborated from eight independent assessment sites.
- Labels cover multiple dimensions of veracity including reliability, bias, transparency, and trust.
- 40 of the 194 sources remained unlabelled after cross-site corroboration.
- All articles are English-language and collected directly from source websites, independent of social media engagement.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.