Skip to main content
QUICK REVIEW

[Paper Review] What has ChatGPT read? The origins of archaeological citations used by a generative artificial intelligence application

Dirk Spennemann|arXiv (Cornell University)|Aug 7, 2023
Topic Modeling4 citations
TL;DR

This study investigates the sources of archaeological citations generated by ChatGPT using cloze analysis to assess whether the model accessed genuine academic texts during training. It finds that while ChatGPT provides plausible references, most are fictitious; genuine citations are likely sourced from Wikipedia rather than original texts, raising concerns about data quality and reliability in AI-generated academic content.

ABSTRACT

The public release of ChatGPT has resulted in considerable publicity and has led to wide-spread discussion of the usefulness and capabilities of generative AI language models. Its ability to extract and summarise data from textual sources and present them as human-like contextual responses makes it an eminently suitable tool to answer questions users might ask. This paper tested what archaeological literature appears to have been included in ChatGPT's training phase. While ChatGPT offered seemingly pertinent references, a large percentage proved to be fictitious. Using cloze analysis to make inferences on the sources 'memorised' by a generative AI model, this paper was unable to prove that ChatGPT had access to the full texts of the genuine references. It can be shown that all references provided by ChatGPT that were found to be genuine have also been cited on Wikipedia pages. This strongly indicates that the source base for at least some of the data is found in those pages. The implications of this in relation to data quality are discussed.

Motivation & Objective

  • To determine the provenance of archaeological references cited by ChatGPT.
  • To assess whether ChatGPT had access to full-text versions of genuine academic sources during its training phase.
  • To evaluate the reliability of AI-generated citations in academic and cultural heritage contexts.
  • To investigate whether Wikipedia serves as a primary source for ChatGPT’s training data in archaeology and cultural heritage.

Proposed method

  • Conducted cloze analysis to infer which sources were 'memorized' by ChatGPT based on its ability to complete reference fragments.
  • Collected and cross-referenced ChatGPT-generated citations against known academic literature in archaeology and cultural heritage management.
  • Verified authenticity of references by checking their presence in academic databases (e.g., JSTOR), Google Books, and Archive.org.
  • Compared citation sources with Wikipedia page citations to identify potential data lineage.
  • Performed repeated queries across multiple topics and model versions to test consistency and reliability of responses.
  • Analyzed citation accuracy by assessing correct author, title, year, and source integrity across 560 total references.

Experimental results

Research questions

  • RQ1What proportion of archaeological citations generated by ChatGPT are factually accurate versus fictitious?
  • RQ2Which sources are most likely to have contributed to ChatGPT’s training data for archaeological references?
  • RQ3Do genuine references cited by ChatGPT appear in Wikipedia, suggesting it may have learned from Wikipedia rather than original texts?
  • RQ4To what extent does ChatGPT access full-text versions of cited academic works during its training?
  • RQ5How consistent are ChatGPT’s citations across repeated queries and model versions?

Key findings

  • Of 560 total references analyzed, 48.8% were confabulated or fictitious, indicating a high rate of hallucination in AI-generated citations.
  • Only 32.1% of references were correctly cited with accurate author, title, and year information.
  • All genuine references cited by ChatGPT were found to be cited on corresponding Wikipedia pages, suggesting Wikipedia as a likely training data source.
  • No evidence was found that ChatGPT accessed full-text versions of the genuine academic references during training.
  • In the Australian archaeology domain, 84.5% of citations were confabulated, indicating a particularly high error rate in this field.
  • The model showed consistent citation errors across repeated queries, with regenerated responses often producing different and equally inaccurate references.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.