Skip to main content
QUICK REVIEW

[论文解读] A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI

Seliem El-Sayed, Canfer Akbulut|arXiv (Cornell University)|Apr 23, 2024
Mental Health Research Topics被引用 7
一句话总结

本文定义了 AI 劝说,区分理性劝说与操控,绘制了伤害类型及其潜在机制,并讨论了面向机制的缓解措施,针对文本生成式 AI 的过程性伤害。

ABSTRACT

Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generative AI presents a new risk profile of persuasion due the opportunity for reciprocal exchange and prolonged interactions. This has led to growing concerns about harms from AI persuasion and how they can be mitigated, highlighting the need for a systematic study of AI persuasion. The current definitions of AI persuasion are unclear and related harms are insufficiently studied. Existing harm mitigation approaches prioritise harms from the outcome of persuasion over harms from the process of persuasion. In this paper, we lay the groundwork for the systematic study of AI persuasion. We first put forward definitions of persuasive generative AI. We distinguish between rationally persuasive generative AI, which relies on providing relevant facts, sound reasoning, or other forms of trustworthy evidence, and manipulative generative AI, which relies on taking advantage of cognitive biases and heuristics or misrepresenting information. We also put forward a map of harms from AI persuasion, including definitions and examples of economic, physical, environmental, psychological, sociocultural, political, privacy, and autonomy harm. We then introduce a map of mechanisms that contribute to harmful persuasion. Lastly, we provide an overview of approaches that can be used to mitigate against process harms of persuasion, including prompt engineering for manipulation classification and red teaming. Future work will operationalise these mitigations and study the interaction between different types of mechanisms of persuasion.

研究动机与目标

  • 定义具说服力的生成式 AI 并区分理性劝说与操纵。
  • 映射 AI 劝说在各领域(经济、心理、政治等)引发的伤害。
  • 识别使具说服力的 AI 的机制与模型特征,以 informing targeted mitigations。
  • 优先考虑过程性伤害并提出缓解方法,如提示工程和红队测试。
  • 为缓解措施的落地化和机制交互研究奠定基础。

提出的方法

  • 为理性说服性和操纵性生成式 AI 输出提出明确定义。
  • 建立 AI 劝说造成的伤害地图,其中包括过程性和结果性伤害(Appendix A)。
  • 提出一个机制为基础的框架,将模型特征与说服能力联系起来(Table 3)。
  • 区分聚焦于过程性伤害而非结果性伤害,以实现可行的缓解。
  • 调查和讨论缓解策略:提示工程、分类、针对说服性机制的分类器、RLHF、可扩展的监督与解释性。
  • 概述未来将缓解措施落地和评估的步骤。
Figure 1: Forms of influence
Figure 1: Forms of influence

实验结果

研究问题

  • RQ1什么构成 AI 劝说及其相关现象?
  • RQ2AI 系统如何劝说,劝说会带来哪些伤害?
  • RQ3哪些机制使 AI 更具说服力,哪些模型特征为其所贡献?
  • RQ4我们如何缓解 AI 劝说的过程性伤害,以及这些与不同情境如何交互?

主要发现

  • 一个基础性的区分被提出:理性劝说(事实和合理推理)与操纵(利用偏见或歪曲信息)。
  • 伤害分为过程性伤害和结果性伤害,并对潜在伤害在各领域的映射(Appendix A)进行了详细说明。
  • 机制映射将模型特征与说服机制联系起来(例如信任/融洽、拟人化、个性化、欺骗、操纵性策略、选择环境改变)。
  • 由于可行性、共识和降低下游伤害的潜力,优先缓解过程性伤害。
  • 缓解策略包括用于分类的提示工程、情境红队化,以及针对有害说服机制的分类器和监督方法的开发。
  • 该工作提供了框架和附录,用于操作化缓解措施并研究说服机制之间的相互作用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。