Skip to main content
QUICK REVIEW

[Paper Review] Formats over Time: Exploring UK Web History

Andrew Jackson|arXiv (Cornell University)|Oct 5, 2012
Web Data Mining and Analysis1 references6 citations
TL;DR

This study analyzes over 2.5 billion UK web resources from 1996 to 2010 using DROID-B and Apache Tika to identify formats, versions, software, and hardware, revealing that software obsolescence is rare due to stabilizing network effects. The key finding is that most formats persist for over a decade, with new formats emerging slowly and versions fading gradually, challenging the notion that digital formats are inherently brittle.

ABSTRACT

Is software obsolescence a significant risk? To explore this issue, we analysed a corpus of over 2.5 billion resources corresponding to the UK Web domain, as crawled between 1996 and 2010. Using the DROID and Apache Tika identification tools, we examined each resource and captured the results as extended MIME types, embedding version, software and hardware identifiers alongside the format information. The combined results form a detailed temporal format profile of the corpus, which we have made available as open data. We present the results of our initial analysis of this dataset. We look at image, HTML and PDF resources in some detail, showing how the usage of different formats, versions and software implementations has changed over time. Furthermore, we show that software obsolescence is rare on the web and uncover evidence indicating that network effects act to stabilise formats against obsolescence.

Motivation & Objective

  • To investigate whether software obsolescence is a significant risk for long-term digital preservation.
  • To understand how format usage, versions, and software implementations evolve over time in the UK web domain.
  • To evaluate the role of network effects in stabilizing digital formats against obsolescence.
  • To create and validate a large-scale, open-format profile of web resources using automated identification tools.
  • To assess the reliability and consistency of format identification tools (DROID-B and Apache Tika) across a massive corpus.

Proposed method

  • Collected a 35TB corpus of 2.5 billion UK web resources from 1996 to 2010, hosted on a 50-node HDFS cluster.
  • Used DROID-B, a modified version of DROID, to identify formats with version, software, and hardware metadata.
  • Complemented with Apache Tika for broader format coverage and deeper bitstream parsing.
  • Performed format identification directly on bitstreams, not metadata or URLs, to ensure content-based accuracy.
  • Combined results from both tools to detect inconsistencies, improve signature quality, and refine format identification.
  • Analyzed temporal trends in format popularity, version usage, and software/hardware identifiers across image, HTML, and PDF resources.

Experimental results

Research questions

  • RQ1To what extent do network effects stabilize digital formats and reduce the risk of obsolescence?
  • RQ2How do the usage patterns of different formats, versions, and software implementations change over time?
  • RQ3Are new formats emerging at a sustainable rate, and are older formats fading quickly or gradually?
  • RQ4How consistent are format identification results between DROID-B and Apache Tika across a large, diverse corpus?
  • RQ5Can the combined output of multiple tools provide a more reliable and granular profile of digital format evolution?

Key findings

  • Forty-seven formats have persisted for 15 years or more, indicating long-term stability rather than rapid obsolescence.
  • On average, about six new formats emerged per year between 1996 and 2010, while only about two became obsolete annually.
  • JPEG remained consistently dominant among image formats, while GIF, TIFF, and XBM saw declining usage, with XBM usage dropping sharply.
  • HTML versions grew gradually over time, with all six major versions active by 2010, and newer versions dominating for several years before fading slowly.
  • PDF versions increased from three (1.0–1.2) in 1996 to eleven by 2010, with 1.2–1.6 remaining dominant, showing a slow decline of older versions.
  • Over 2,100 distinct software implementations were identified for PDF, and over 1,900 for JPEG, suggesting high standardization and maturity in widely used formats.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.