Skip to main content
QUICK REVIEW

[Paper Review] Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations

Josh A. Goldstein, Girish Sastry|arXiv (Cornell University)|Jan 10, 2023
Topic Modeling146 citations
TL;DR

The paper assesses how generative language models could transform influence operations and surveys mitigations using a kill-chain framework, emphasizing that no silver bullet exists and that collective action is needed.

ABSTRACT

Generative language models have improved drastically, and can now produce realistic text outputs that are difficult to distinguish from human-written content. For malicious actors, these language models bring the promise of automating the creation of convincing and misleading text for use in influence operations. This report assesses how language models might change influence operations in the future, and what steps can be taken to mitigate this threat. We lay out possible changes to the actors, behaviors, and content of online influence operations, and provide a framework for stages of the language model-to-influence operations pipeline that mitigations could target (model construction, model access, content dissemination, and belief formation). While no reasonable mitigation can be expected to fully prevent the threat of AI-enabled influence operations, a combination of multiple mitigations may make an important difference.

Motivation & Objective

  • Assess how language models could alter actors, behaviors, and content in influence operations.
  • Survey potential threats and a range of mitigations across the influence campaign pipeline.
  • Highlight critical unknowns and the need for coordinated policy, technical, and societal responses.
  • Provide a framework for evaluating mitigations and identify research directions.

Proposed method

  • Literature review and synthesis of existing disinformation research.
  • Workshop-informed analysis from a multidisciplinary team.
  • Applying the ABCs of disinformation (Actors, Behaviors, Content) to language models.
  • Developing a kill-chain based mitigation framework spanning model construction, access, content dissemination, and belief formation.
  • Contextual discussion of generative model progress and access diffusion to ground threat assessment.

Experimental results

Research questions

  • RQ1How might language models change the actors that wage influence operations?
  • RQ2How might language models alter the behaviors and tactics used in influence operations?
  • RQ3How might language models affect the content produced in influence campaigns and its impact?
  • RQ4What mitigations could effectively reduce the impact of AI-enabled influence operations across the pipeline?
  • RQ5What governance and collaboration mechanisms are needed to implement these mitigations?

Key findings

  • Language models are likely to be useful for influence operations in the future due to improvements in usability, reliability, and efficiency.
  • There is no single mitigation that fully prevents AI-enabled influence operations; a whole-of-society approach is required.
  • Effective mitigation will require cooperation among AI developers, social platforms, governments, and civil society across multiple stages of the operation pipeline.
  • Mitigations should address both supply (model design, access) and demand/dissemination (content provenance, media literacy, platform measures).
  • The most radical options (e.g., internet provenance standards) would require extreme coordination and may not be desirable; many mitigations require further development and scrutiny.
  • The report emphasizes evaluating mitigations with a framework and identifying avenues for further research.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.