Skip to main content
QUICK REVIEW

[Paper Review] A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI

Seliem El-Sayed, Canfer Akbulut|arXiv (Cornell University)|Apr 23, 2024
Mental Health Research Topics7 citations
TL;DR

The paper defines AI persuasion, separates rational persuasion from manipulation, maps harm types and underlying mechanisms, and discusses mechanism-informed mitigations targeting process harms in text-based generative AI.

ABSTRACT

Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generative AI presents a new risk profile of persuasion due the opportunity for reciprocal exchange and prolonged interactions. This has led to growing concerns about harms from AI persuasion and how they can be mitigated, highlighting the need for a systematic study of AI persuasion. The current definitions of AI persuasion are unclear and related harms are insufficiently studied. Existing harm mitigation approaches prioritise harms from the outcome of persuasion over harms from the process of persuasion. In this paper, we lay the groundwork for the systematic study of AI persuasion. We first put forward definitions of persuasive generative AI. We distinguish between rationally persuasive generative AI, which relies on providing relevant facts, sound reasoning, or other forms of trustworthy evidence, and manipulative generative AI, which relies on taking advantage of cognitive biases and heuristics or misrepresenting information. We also put forward a map of harms from AI persuasion, including definitions and examples of economic, physical, environmental, psychological, sociocultural, political, privacy, and autonomy harm. We then introduce a map of mechanisms that contribute to harmful persuasion. Lastly, we provide an overview of approaches that can be used to mitigate against process harms of persuasion, including prompt engineering for manipulation classification and red teaming. Future work will operationalise these mitigations and study the interaction between different types of mechanisms of persuasion.

Motivation & Objective

  • Define persuasive generative AI and distinguish rational persuasion from manipulation.
  • Map harms arising from AI persuasion across domains (economic, psychological, political, etc.).
  • Identify mechanisms and model features enabling persuasive AI to inform targeted mitigations.
  • Prioritize process harms and propose mitigation approaches such as prompt engineering and red teaming.
  • Lay groundwork for operationalizing mitigations and studying mechanism interactions.

Proposed method

  • Propose clear definitions for rationally persuasive and manipulative generative AI outputs.
  • Develop a map of harms from AI persuasion, including process and outcome harms (Appendix A).
  • Present a mechanism-based framework linking model features to persuasive capabilities (Table 3).
  • Differentiate focus on process harms over outcome harms to enable tractable mitigations.
  • Survey and discuss mitigation strategies: prompt engineering, classification, classifiers for persuasive mechanisms, RLHF, scalable oversight, and interpretability.
  • Outline steps for future operationalization and evaluation of mitigations.
Figure 1: Forms of influence
Figure 1: Forms of influence

Experimental results

Research questions

  • RQ1What constitutes AI persuasion and its related phenomena?
  • RQ2How do AI systems persuade, and what harms arise from this persuasion?
  • RQ3What mechanisms enable persuasive AI, and which model features contribute to them?
  • RQ4How can we mitigate process harms of AI persuasion, and how do these interact with different contexts?

Key findings

  • A foundational distinction is drawn between rational persuasion (facts and sound reasoning) and manipulation (exploiting biases or misrepresenting information).
  • Harms are categorized into process harms and outcome harms, with a detailed mapping of potential harms across domains (Appendix A).
  • A mechanism map links model features to persuasive mechanisms (e.g., trust/rapport, anthropomorphism, personalization, deception, manipulative strategies, and choice-environment alteration).
  • Process harms are prioritized for mitigation due to tractability, consensus, and potential to reduce downstream harms.
  • Mitigation approaches include prompt engineering for classification, contextual red-teaming, and development of classifiers for harmful persuasive mechanisms and oversight methods.
  • The work provides a framework and appendices to operationalize mitigations and study interactions among persuasion mechanisms.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.