Skip to main content
QUICK REVIEW

[Paper Review] Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus

Jack Bandy, Nicholas Vincent|arXiv (Cornell University)|May 11, 2021
Topic Modeling18 references33 citations
TL;DR

The paper applies the datasheet framework to BookCorpus to document its motivation, composition, collection, and potential deficiencies, highlighting copyright, duplication, and genre-skew concerns.

ABSTRACT

Recent literature has underscored the importance of dataset documentation work for machine learning, and part of this work involves addressing "documentation debt" for datasets that have been used widely but documented sparsely. This paper aims to help address documentation debt for BookCorpus, a popular text dataset for training large language models. Notably, researchers have used BookCorpus to train OpenAI's GPT-N models and Google's BERT models, even though little to no documentation exists about the dataset's motivation, composition, collection process, etc. We offer a preliminary datasheet that provides key context and information about BookCorpus, highlighting several notable deficiencies. In particular, we find evidence that (1) BookCorpus likely violates copyright restrictions for many books, (2) BookCorpus contains thousands of duplicated books, and (3) BookCorpus exhibits significant skews in genre representation. We also find hints of other potential deficiencies that call for future research, including problematic content, potential skews in religious representation, and lopsided author contributions. While more work remains, this initial effort to provide a datasheet for BookCorpus adds to growing literature that urges more careful and systematic documentation for machine learning datasets.

Motivation & Objective

  • Motivate the need for dataset documentation in ML research (documentation debt).
  • Provide a structured datasheet for BookCorpus to capture motivation, composition, collection, and usage considerations.
  • Identify key deficiencies and potential ethical and legal risks in BookCorpus to guide future use.
  • Offer recommendations for better documentation practices and future research directions in dataset governance.

Proposed method

  • Apply the datasheet framework (Gebru et al.) to BookCorpus, including questions on motivation, composition, collection, cleaning, uses, and distribution.
  • Collect and compare three versions of BookCorpus: the original 2014 BookCorpus, BookCorpusOpen (2020/2021), and Smashwords21 metadata.
  • Systematically analyze the dataset for copyright issues, duplicates, and skews across genres and religious representation.
  • Document collection processes, licensing, consent, and potential impact on data subjects.

Experimental results

Research questions

  • RQ1What were the original motivations and use cases for BookCorpus, and who funded its creation?
  • RQ2What is the composition of BookCorpus in terms of books, words, and genres, and how does it vary across versions?
  • RQ3What are the potential deficiencies and risks (copyright, duplicates, content sensitivity, sampling bias) present in BookCorpus?
  • RQ4How was BookCorpus collected, cleaned, distributed, and maintained, and what ethical considerations apply?
  • RQ5What implications do these findings have for using BookCorpus in current and future ML research?

Key findings

  • Only 7,185 unique books exist in BookCorpus, with 2,930 duplicates identified.
  • BookCorpus likely violates copyright restrictions for many books, based on observed licensing statements.
  • Significant genre skew exists, with Romance substantially over-represented compared to newer copies and the Smashwords21 superset.
  • Presence of potentially problematic content and skewed religious representation noted, requiring caution.
  • Personal contact information (email addresses) found in the data, indicating sensitivity and privacy considerations.
  • BookCorpus is not publicly maintained; multiple replication versions exist (BookCorpusOpen, Smashwords21) and access is fragmented.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.