[논문 리뷰] A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
본 논문은 AI 설득을 정의하고 합리적 설득과 조작을 구분하며, 피해 유형과 기저 메커니즘을 매핑하고, 텍스트 기반 생성 AI의 과정상 해를 대상으로 한 메커니즘 기반 완화책을 논의한다.
Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generative AI presents a new risk profile of persuasion due the opportunity for reciprocal exchange and prolonged interactions. This has led to growing concerns about harms from AI persuasion and how they can be mitigated, highlighting the need for a systematic study of AI persuasion. The current definitions of AI persuasion are unclear and related harms are insufficiently studied. Existing harm mitigation approaches prioritise harms from the outcome of persuasion over harms from the process of persuasion. In this paper, we lay the groundwork for the systematic study of AI persuasion. We first put forward definitions of persuasive generative AI. We distinguish between rationally persuasive generative AI, which relies on providing relevant facts, sound reasoning, or other forms of trustworthy evidence, and manipulative generative AI, which relies on taking advantage of cognitive biases and heuristics or misrepresenting information. We also put forward a map of harms from AI persuasion, including definitions and examples of economic, physical, environmental, psychological, sociocultural, political, privacy, and autonomy harm. We then introduce a map of mechanisms that contribute to harmful persuasion. Lastly, we provide an overview of approaches that can be used to mitigate against process harms of persuasion, including prompt engineering for manipulation classification and red teaming. Future work will operationalise these mitigations and study the interaction between different types of mechanisms of persuasion.
연구 동기 및 목표
- 설득적 생성 AI를 정의하고 합리적 설득과 조작을 구분한다.
- AI 설득으로부터 발생하는 피해를 경제적, 심리적, 정치적 등 다양한 영역에서 매핑한다.
- 설득 가능성을 야기하는 메커니즘과 모델 특성을 식별하여 타깃화된 완화를 위한 정보를 제공한다.
- 과정상의 피해에 우선순위를 두고 프롬프트 엔지니어링과 레드팀핑(적대적 테스트) 같은 완화 방식을 우선 제시한다.
- 완화책의 실행 가능화를 위한 기반을 마련하고 메커니즘 간 상호작용을 연구한다.
제안 방법
- 합리적으로 설득적인 출력과 조작적 출력에 대한 명확한 정의를 제안한다.
- 과정상 피해와 결과 피해를 포함한 AI 설득으로부터의 피해 맵을 개발한다 (부록 A).
- 모델 특징을 설득적 능력에 연결하는 기계 기반 프레임워크를 제시한다 (표 3).
- 실행 가능하고 관리 가능한 완화를 가능케 하기 위해 과정 피해에 집중하고 결과 피해를 구분한다.
- 프롬프트 엔지니어링, 분류, 설득 메커니즘용 분류기, RLHF, 확장 가능한 감독, 해석 가능성 등 완화 전략을 조사하고 논의한다.
- 완화책의 향후 실행화와 평가를 위한 단계들을 개괄한다.

실험 결과
연구 질문
- RQ1AI 설득과 그에 관련된 현상은 무엇인가?
- RQ2AI 시스템은 어떻게 설득하며 이 설득에서 어떤 피해가 발생하는가?
- RQ3설득 AI를 가능하게 하는 메커니즘은 무엇이며, 어떤 모델 특성들이 그것에 기여하는가?
- RQ4AI 설득의 과정 피해를 어떻게 완화할 수 있으며, 이것이 다양한 맥락과 어떻게 상호작용하는가?
주요 결과
- 합리적 설득(사실과 타당한 추론)과 조작(편향을 악용하거나 정보를 왜곡하는 행위) 사이의 기본적 차이가 제시된다.
- 피해는 과정 피해와 결과 피해로 분류되며, 다양한 영역에서의 잠재적 피해를 상세히 매핑한다(부록 A).
- 메커니즘 맵은 모델 특징을 설득 메커니즘(예: 신뢰/라포, 의인화, 개인화, 기만, 조작적 전략, 선택-환경 변화)에 연결한다.
- 과정 피해를 다루기 쉬움, 합의 형성, 그리고 하류 피해 감소 가능성으로 인해 완화 우선순위가 부여된다.
- 완화 전략에는 분류를 위한 프롬프트 엔지니어링, 맥락적 레드팀링, 해로운 설득 메커니즘 분류기 개발 및 감독 방법이 포함된다.
- 본 연구는 완화를 실행화하고 설득 메커니즘 간의 상호작용을 연구하기 위한 프레임워크와 부록을 제공한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.