Skip to main content
QUICK REVIEW

[Paper Review] The Use of Web Archives in Disinformation Research

Michele C. Weigle|arXiv (Cornell University)|Jun 16, 2023
Misinformation and Its ImpactsSocial Sciences3 citations
TL;DR

This paper demonstrates how web archives—particularly the Internet Archive’s Wayback Machine—are critical tools for disinformation research, enabling verification of deleted content, analysis of social media changes, and tracking of disinformation campaigns. By leveraging archived mementos, researchers can detect webpage edits, validate historical claims, and study metadata of suspended accounts, significantly enhancing transparency and accountability in digital discourse.

ABSTRACT

In recent years, journalists and other researchers have used web archives as an important resource for their study of disinformation. This paper provides several examples of this use and also brings together some of the work that the Old Dominion University Web Science and Digital Libraries (WS-DL) research group has done in this area. We will show how web archives have been used to investigate changes to webpages, study archived social media including deleted content, and study known disinformation that has been archived.

Motivation & Objective

  • To investigate how web archives support disinformation research by preserving historical web content that may be deleted or altered.
  • To address the challenge of verifying the authenticity of screenshots and social media posts by using archived web versions as evidence.
  • To develop tools and workflows that enable journalists and researchers to systematically analyze changes in webpages and social media content over time.
  • To explore the role of web archives in preserving and analyzing disinformation, including archived advertisements and metadata from suspended accounts.
  • To promote the use of web archives as reliable, tamper-resistant sources for digital forensics and media literacy research.

Proposed method

  • Utilizing the Wayback Machine’s memento storage to compare historical and current versions of webpages for content changes.
  • Applying the MemGator tool and Memento Time Travel service to query and retrieve archived web content from multiple independent web archives.
  • Developing a prototype interface to search for deleted terms and phrases across archived webpages, with support for animated diff visualizations.
  • Using the Wayback Machine’s CDX API to retrieve all archived tweets for a specific user by URL pattern matching (e.g., /realdonaldtrump/status/*).
  • Analyzing metadata from archived social media profiles and posts to study follower growth, engagement patterns, and network structures of disinformation actors.
  • Leveraging web archives to study archived advertisements and their role in shaping public perception during events like the COVID-19 pandemic.

Experimental results

Research questions

  • RQ1How can web archives be used to detect and verify changes to political webpages, such as the removal of controversial statements?
  • RQ2To what extent can archived web content serve as reliable evidence to corroborate or refute fabricated social media screenshots?
  • RQ3How do web archives support the analysis of disinformation campaigns by preserving deleted or suspended social media content?
  • RQ4What role do archived advertisements play in understanding the evolution of disinformation strategies over time?
  • RQ5How can tools built on web archive data improve the detection and tracking of disinformation actors and their networks?

Key findings

  • Journalists successfully used the Wayback Machine to document that multiple 2022 U.S. Senate candidates, including Blake Masters and George Santos, removed controversial statements from their campaign websites.
  • A prototype tool developed by Frew et al. (2023) enables users to search for deleted phrases across archived webpages and visualize content changes via animated diffs or sliding interfaces.
  • The MemGator tool and Memento Time Travel service allowed researchers to verify Joy Ann Reid’s controversial blog posts by retrieving mementos from multiple independent archives, countering claims of image manipulation.
  • The Internet Archive began labeling known disinformation in web archives, such as a Medium article removed for policy violations but still shared via Wayback Machine links.
  • Archived Twitter data via the CDX interface revealed that over 100,000 tweets from @realdonaldtrump were preserved in the Wayback Machine, enabling longitudinal analysis of his social media presence.
  • Research on archived advertisements, including those promoting COVID-19 masks, revealed that such content is critical for understanding the historical context of public health messaging and disinformation campaigns.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.