Skip to main content
QUICK REVIEW

[Paper Review] AI and the Problem of Knowledge Collapse

A. Townsend Peterson|arXiv (Cornell University)|Apr 4, 2024
Big Data and Business Intelligence4 citations
TL;DR

This paper introduces the concept of 'knowledge collapse'—a societal shift toward narrow, AI-mediated understanding due to overreliance on large language models (LLMs) that favor central, common responses over diverse perspectives. Using a simulation model, the study shows that a 20% cost discount on AI-generated content can push public beliefs 2.3 times further from the truth than when relying on full-spectrum knowledge, highlighting risks to innovation and cultural richness.

ABSTRACT

While artificial intelligence has the potential to process vast amounts of data, generate new insights, and unlock greater productivity, its widespread adoption may entail unforeseen consequences. We identify conditions under which AI, by reducing the cost of access to certain modes of knowledge, can paradoxically harm public understanding. While large language models are trained on vast amounts of diverse data, they naturally generate output towards the 'center' of the distribution. This is generally useful, but widespread reliance on recursive AI systems could lead to a process we define as "knowledge collapse", and argue this could harm innovation and the richness of human understanding and culture. However, unlike AI models that cannot choose what data they are trained on, humans may strategically seek out diverse forms of knowledge if they perceive them to be worthwhile. To investigate this, we provide a simple model in which a community of learners or innovators choose to use traditional methods or to rely on a discounted AI-assisted process and identify conditions under which knowledge collapse occurs. In our default model, a 20% discount on AI-generated content generates public beliefs 2.3 times further from the truth than when there is no discount. An empirical approach to measuring the distribution of LLM outputs is provided in theoretical terms and illustrated through a specific example comparing the diversity of outputs across different models and prompting styles. Finally, based on the results, we consider further research directions to counteract such outcomes.

Motivation & Objective

  • To investigate how widespread adoption of AI-assisted knowledge access may paradoxically reduce public understanding by narrowing the diversity of knowledge sources.
  • To model the conditions under which individuals or communities may strategically seek out diverse knowledge despite AI’s cost discounting.
  • To empirically measure the diversity of LLM outputs across models and prompting strategies, particularly focusing on long-tail knowledge.
  • To identify systemic risks of knowledge collapse in education, innovation, and cultural preservation due to recursive AI mediation of information.
  • To propose research directions for mitigating knowledge collapse through improved AI design, transparency, and user incentives for diverse knowledge sourcing.

Proposed method

  • Develops a positive knowledge spillovers model where agents choose between using AI-assisted content (with a cost discount) or investing in full-spectrum, diverse knowledge sources.
  • Uses a simulation framework to model belief formation in a community, tracking how AI cost discounts affect the distance of public beliefs from the true distribution of knowledge.
  • Employs a theoretical framework to define and measure output diversity in LLMs using frequency distributions of named entities (e.g., philosophers, schools of thought) across different prompts.
  • Compares LLM outputs across models (GPT-3.5-turbo, Claude-3-sonnet, Gemini-Pro, Llama2-70b) under varying prompting styles—simple queries vs. structured requests for diverse responses.
  • Truncates frequency distributions to the 600 most frequent entities to visualize and compare the concentration of responses across models and prompts.
  • Uses empirical data from 2,693 identified philosophical entities to assess how prompting style influences the representation of long-tail knowledge.

Experimental results

Research questions

  • RQ1Under what conditions does AI-mediated access to knowledge lead to a systematic divergence of public beliefs from the true distribution of knowledge?
  • RQ2How does a cost discount on AI-generated content affect the diversity and representativeness of knowledge accessed by individuals and communities?
  • RQ3To what extent can human strategic behavior—such as seeking out non-AI sources—prevent knowledge collapse in the face of AI’s efficiency advantage?
  • RQ4How do different prompting strategies influence the diversity of LLM outputs, particularly in representing underrepresented or long-tail knowledge?
  • RQ5What are the implications of knowledge collapse for innovation, cultural preservation, and equitable access to diverse worldviews?

Key findings

  • A 20% cost discount on AI-generated content increases the distance of public beliefs from the truth by 2.3 times compared to a scenario with no discount.
  • LLMs systematically favor central, high-frequency responses, leading to a concentration of output around dominant perspectives and a neglect of long-tail knowledge.
  • Prompting strategies that explicitly request diverse responses (e.g., listing 20 philosophers from a specific region) significantly increase the representation of underrepresented entities, reducing dominance of a few names.
  • Even with high-quality models like GPT-3.5-turbo and Claude-3-sonnet, the default output distribution remains highly skewed, with a few entities (e.g., Aristotle, Kant) dominating frequency counts.
  • The frequency of mentions for less common entities—such as Yoruba philosophy, Hausa philosophy, and Avicenna—remains low across models, indicating persistent marginalization in AI outputs.
  • When prompts explicitly encourage diversity (e.g., prompt v5), no single entity dominates, demonstrating that output diversity can be significantly improved through intentional prompting design.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.