Skip to main content
QUICK REVIEW

[Paper Review] Sociotechnical Implications of Generative Artificial Intelligence for Information Access

Bhaskar Mitra, Henriette Cramer|arXiv (Cornell University)|May 19, 2024
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This paper examines the sociotechnical implications of generative AI in information access, using a consequences-mechanisms-risks (CMR) framework to identify systemic risks such as information ecosystem disruption, power concentration, bias amplification, and environmental harm. It proposes evaluation and mitigation strategies, emphasizing ethical design, data labor advocacy, and re-centering societal needs in AI-driven information retrieval systems.

ABSTRACT

Robust access to trustworthy information is a critical need for society with implications for knowledge production, public health education, and promoting informed citizenry in democratic societies. Generative AI technologies may enable new ways to access information and improve effectiveness of existing information retrieval systems but we are only starting to understand and grapple with their long-term social implications. In this chapter, we present an overview of some of the systemic consequences and risks of employing generative AI in the context of information access. We also provide recommendations for evaluation and mitigation, and discuss challenges for future research.

Motivation & Objective

  • To analyze the systemic sociotechnical risks of generative AI in information access, particularly focusing on large language models (LLMs).
  • To identify how generative AI disrupts information ecosystems, concentrates power, and exacerbates inequities in knowledge production and access.
  • To examine the role of data labor, environmental costs, and AI alignment in shaping ethical and sustainable information retrieval systems.
  • To propose actionable evaluation and mitigation strategies grounded in sociotechnical systems thinking and stakeholder-centered design.
  • To challenge the current trajectory of AI in IR by re-centering societal needs and envisioning emancipatory futures for information access.

Proposed method

  • Adopts the consequences-mechanisms-risks (CMR) framework from Gausen et al. to structure analysis of sociotechnical impacts.
  • Maps identified consequences (e.g., information ecosystem disruption) to underlying mechanisms (e.g., content pollution, direct model access) and associated risks.
  • Synthesizes existing literature on LLM risks, including bias, alignment, data labor, and environmental impact, within the context of information retrieval.
  • Proposes data-focused mitigation strategies such as data dividends, collective action (e.g., data strikes), and alternative licensing models.
  • Integrates insights from labor organizing (e.g., Hollywood strikes) and technical countermeasures (e.g., data poisoning tools like NightShade, Glaze, Mist).
  • Reframes IR research through Mitra’s hierarchy of stakeholder needs, advocating for a shift toward societal well-being and equity in system design.

Experimental results

Research questions

  • RQ1How do generative AI systems disrupt information ecosystems, and what mechanisms underlie these disruptions?
  • RQ2What are the systemic risks of power concentration, bias amplification, and environmental degradation in AI-driven information access?
  • RQ3How can data labor be ethically recognized and compensated in the training of generative AI models?
  • RQ4What role can collective action and alternative business models play in mitigating harms from generative AI in information retrieval?
  • RQ5How can IR research and system development be re-centered around societal needs rather than technological feasibility alone?

Key findings

  • Generative AI introduces systemic risks such as information ecosystem disruption through content pollution, search engine manipulation, and the 'game of telephone' effect in knowledge propagation.
  • Concentration of power in large AI developers is driven by compute and data moats, enabling AI persuasion and alignment challenges that threaten democratic and equitable information access.
  • Marginalization of underrepresented voices and bias amplification are exacerbated by the appropriation of data labor and lack of representation in training data and model development.
  • Environmental costs of training and deploying LLMs, including high resource demand and waste, pose significant long-term risks to sustainability and ecological equity.
  • Mitigation strategies such as data dividends, data strikes, and collective organizing (e.g., inspired by Hollywood strikes) show promise in empowering data producers and reasserting stakeholder control.
  • A fundamental shift in IR research is needed—re-centering societal needs through frameworks like Mitra’s hierarchy of stakeholder needs to envision emancipatory, equitable futures for information access.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.