Skip to main content
QUICK REVIEW

[Paper Review] NELA-GT-2018: A Large Multi-Labelled News Dataset for The Study of Misinformation in News Articles

Jeppe Nørregaard, Benjamin D. Horne|arXiv (Cornell University)|Apr 2, 2019
Misinformation and Its Impacts15 references38 citations
TL;DR

Introduces a large, engagement-independent dataset of 713,534 English-language news articles from 194 outlets (02/2018–11/2018) with multi-site ground-truth labels for veracity and credibility.

ABSTRACT

In this paper, we present a dataset of 713k articles collected between 02/2018-11/2018. These articles are collected directly from 194 news and media outlets including mainstream, hyper-partisan, and conspiracy sources. We incorporate ground truth ratings of the sources from 8 different assessment sites covering multiple dimensions of veracity, including reliability, bias, transparency, adherence to journalistic standards, and consumer trust. The NELA-GT-2018 dataset can be found at https://doi.org/10.7910/DVN/ULHLCB.

Motivation & Objective

  • Fill the gap of large, multi-dimensional ground-truth labels for misinformation research independent of social media engagement.
  • Provide a long-span dataset suitable for machine learning and qualitative studies of news veracity and bias.
  • Corroborate source-level ground truth across multiple assessment platforms to enable diverse analyses.

Proposed method

  • Systematically crawled RSS feeds of 194 news outlets twice daily from 02/2018 to 11/2018 to collect 713,534 articles.
  • Compiled ground-truth labels for sources from eight assessment sites covering reliability, bias, transparency, and trust.
  • Organized article data in an SQLite database with date, source, title, and cleaned text; labels stored in CSV per source.
  • Compared and described the eight ground-truth sources (NewsGuard, Pew, Wikipedia, OpenSources, MBFC, AllSides, BuzzFeed News, PolitiFact) and their labeling schemes.
  • Provided a documented dataset with accompanying use-case guidance for distant supervision and semi-supervised learning.
  • Proposed that the dataset supports ground-truth-driven machine learning and mixed-method research into misinformation tactics.

Experimental results

Research questions

  • RQ1How can a large, multi-label ground-truth dataset enable robust misinformation research independent of engagement signals?
  • RQ2What are the characteristics and coverage of veracity and bias across a diverse set of news sources over an extended time period?
  • RQ3How can source-level ground truth from multiple assessment sites be leveraged for article-level analyses and model training?

Key findings

  • The dataset contains 713,534 articles from 194 news outlets collected between 02/2018 and 11/2018.
  • Ground truth labels are corroborated from eight independent assessment sites.
  • Labels cover multiple dimensions of veracity including reliability, bias, transparency, and trust.
  • 40 of the 194 sources remained unlabelled after cross-site corroboration.
  • All articles are English-language and collected directly from source websites, independent of social media engagement.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.