[Paper Review] Can The Crowd Identify Misinformation Objectively? The Effects of Judgment Scale and Assessor's Background
This study investigates whether non-expert crowd workers can reliably assess the truthfulness of political statements using multiple judgment scales. Despite low inter-annotator agreement, crowdsourced labels—when aggregated across scales—show strong correlation with expert fact-checking assessments, and assessors' implicit political orientation significantly affects judgment accuracy, particularly on polarizing issues like border walls.
Truthfulness judgments are a fundamental step in the process of fighting misinformation, as they are crucial to train and evaluate classifiers that automatically distinguish true and false statements. Usually such judgments are made by experts, like journalists for political statements or medical doctors for medical statements. In this paper, we follow a different approach and rely on (non-expert) crowd workers. This of course leads to the following research question: Can crowdsourcing be reliably used to assess the truthfulness of information and to create large-scale labeled collections for information credibility systems? To address this issue, we present the results of an extensive study based on crowdsourcing: we collect thousands of truthfulness assessments over two datasets, and we compare expert judgments with crowd judgments, expressed on scales with various granularity levels. We also measure the political bias and the cognitive background of the workers, and quantify their effect on the reliability of the data provided by the crowd.
Motivation & Objective
- To evaluate the reliability of non-expert crowd workers in assessing the truthfulness of political statements.
- To examine how different judgment scale granularities affect the quality of truthfulness assessments.
- To investigate the impact of assessors' political background and cognitive abilities on judgment accuracy.
- To compare crowd-generated truthfulness labels with expert fact-checking assessments.
- To understand the sources and strategies used by crowd workers when evaluating misinformation.
Proposed method
- Conducted a large-scale crowdsourcing experiment with US-based workers assessing truthfulness of statements from US and Australian politicians.
- Used a controlled, custom web search engine to standardize information access during the fact-checking task.
- Employed three truthfulness judgment scales: S3 (true/false/neutral), S6 (three-point scale), and S100 (100-point continuous scale).
- Collected data on assessors' political stance (explicit and implicit), cognitive abilities, and geographical relevance to the statements.
- Applied statistical tests (Kruskal-Wallis H, Dunn’s post-hoc) to analyze the impact of political orientation on judgment discernment.
- Aggregated crowd judgments by merging adjacent categories to improve signal-to-noise ratio and assess correlation with expert labels.
Experimental results
Research questions
- RQ1RQ1: Are the used assessment scales suitable to gather, by means of crowdsourcing, truthfulness labels on political statements?
- RQ2RQ2: What is the relationship and agreement between crowd and expert labels, and between labels collected using different scales?
- RQ3RQ3: Which sources of information do crowd workers use to identify online misinformation?
- RQ4RQ4: What is the effect and role of assessors’ background in objectively identifying online misinformation?
Key findings
- Crowd workers showed low inter-annotator agreement across all scales (S3, S6, S100), but aggregated judgments significantly correlated with expert assessments.
- The choice of judgment scale did not affect the overall quality of truthfulness assessments, as all scales performed similarly in capturing true/false distinctions.
- Workers who opposed building a wall along the southern U.S. border demonstrated significantly higher discernment scores (p=0.007) in distinguishing true from false statements.
- Implicit political orientation—particularly on immigration policy—had a statistically significant impact on judgment quality (p=0.004 for PolitiFact statements), while explicit party identification did not.
- Workers tended to rely on sources from the first page of search results, indicating strategic information-seeking behavior during fact-checking.
- No significant difference in judgment quality was observed based on stance on climate change, suggesting issue-specific bias effects.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.