[論文レビュー] A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
本論文はAI説得を定義し、合理的説得と操作の区別を明確にし、害の種類と根本的メカニズムを整理し、テキストベースの生成AIにおけるプロセス害を対象とした機構に基づく緩和策を論じる。
Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generative AI presents a new risk profile of persuasion due the opportunity for reciprocal exchange and prolonged interactions. This has led to growing concerns about harms from AI persuasion and how they can be mitigated, highlighting the need for a systematic study of AI persuasion. The current definitions of AI persuasion are unclear and related harms are insufficiently studied. Existing harm mitigation approaches prioritise harms from the outcome of persuasion over harms from the process of persuasion. In this paper, we lay the groundwork for the systematic study of AI persuasion. We first put forward definitions of persuasive generative AI. We distinguish between rationally persuasive generative AI, which relies on providing relevant facts, sound reasoning, or other forms of trustworthy evidence, and manipulative generative AI, which relies on taking advantage of cognitive biases and heuristics or misrepresenting information. We also put forward a map of harms from AI persuasion, including definitions and examples of economic, physical, environmental, psychological, sociocultural, political, privacy, and autonomy harm. We then introduce a map of mechanisms that contribute to harmful persuasion. Lastly, we provide an overview of approaches that can be used to mitigate against process harms of persuasion, including prompt engineering for manipulation classification and red teaming. Future work will operationalise these mitigations and study the interaction between different types of mechanisms of persuasion.
研究の動機と目的
- 説得的な生成AIを定義し、合理的説得と操作(ミスリードを含む操作)を区別する。
- 経済的・心理的・政治的など、さまざまな領域にわたるAI説得による害を整理する。
- 説得的AIを可能にする機構とモデル特徴を特定し、標的を絞った緩和策の設計に役立てる。
- プロセス害を優先し、プロンプト設計やレッドチーミングなどの緩和アプローチを提案する。
- 緩和策の運用化とメカニズムの相互作用の研究の地盤を築く。
提案手法
- 合理的に説得的な出力と操作的な生成AI出力の明確な定義を提案する。
- AI説得による害のマップを作成し、プロセス害とアウトカム害を含む(Appendix A)。
- モデル特徴を説得能力に結びつける機構ベースの枠組みを提示する(Table 3)。
- 処理可能な緩和を実現するため、アウトカム害よりもプロセス害に焦点を絞ることを区別する。
- 緩和戦略を調査・検討する:プロンプトエンジニアリング、分類、説得機構の分類器、RLHF、スケーラブルな監視、解釈可能性。
- 今後の緩和策の運用化と評価に向けた手順を概説する。

実験結果
リサーチクエスチョン
- RQ1AI説得とそれに関連する現象とは何か?
- RQ2AIシステムはどのように説得するのか、そしてこの説得からどんな害が生じるのか?
- RQ3説得的AIを可能にするメカニズムは何か、それを生み出すモデル特徴はどれか?
- RQ4AI説得のプロセス害をどのように緩和できるか、そしてそれらは異なる文脈とどう相互作用するのか?
主な発見
- 合理的説得(事実と筋の通った推論)と操作(バイアスを利用したり情報を誤って伝えること)との基本的な区別が示されている。
- 害はプロセス害とアウトカム害に分類され、領域別の潜在的な害の詳細なマッピング(Appendix A)を含む。
- メカニズムマップは、モデル特徴を説得機構(例:信頼・ラポール、擬人化、個別化、欺瞞、操作的戦略、選択環境の変更)に結びつける。
- 処理の害は、実行可能性・合意形成・下流の害を低減する潜在力の観点から緩和の優先度が高い。
- 緩和アプローチには、分類のためのプロンプトエンジニアリング、文脈に応じたレッドチーミング、有害な説得機構の分類器の開発、監視手法の開発が含まれる。
- 本研究は、緩和策を運用化し、説得メカニズム間の相互作用を研究するための枠組みと付録を提供する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。