Skip to main content
QUICK REVIEW

[논문 리뷰] Deceptive AI systems that give explanations are more convincing than honest AI systems and can amplify belief in misinformation

Valdemar Danry, Pat Pataranutaporn|arXiv (Cornell University)|2024. 07. 31.
Adversarial Robustness in Machine Learning인용 수 4
한 줄 요약

이 연구는 논리적으로 오류가 있지만 설득력 있는 설명을 생성하는 기만적 AI 시스템이 솔직한 AI 설명이나 단순한 오분류보다 더 설득력이 있으며, 잘못된 정보에 대한 신뢰를 크게 증폭시킨다는 것을 입증한다. 핵심 발견은 설명의 논리적 타당성—즉, 추론이 결론을 뒷받침하는지 여부—가 설명의 설득력에 결정적인 영향을 미친다는 점이며, 타당성이 없는 설명은 더 낮은 신뢰도를 갖는다.

ABSTRACT

Advanced Artificial Intelligence (AI) systems, specifically large language models (LLMs), have the capability to generate not just misinformation, but also deceptive explanations that can justify and propagate false information and erode trust in the truth. We examined the impact of deceptive AI generated explanations on individuals' beliefs in a pre-registered online experiment with 23,840 observations from 1,192 participants. We found that in addition to being more persuasive than accurate and honest explanations, AI-generated deceptive explanations can significantly amplify belief in false news headlines and undermine true ones as compared to AI systems that simply classify the headline incorrectly as being true/false. Moreover, our results show that personal factors such as cognitive reflection and trust in AI do not necessarily protect individuals from these effects caused by deceptive AI generated explanations. Instead, our results show that the logical validity of AI generated deceptive explanations, that is whether the explanation has a causal effect on the truthfulness of the AI's classification, plays a critical role in countering their persuasiveness - with logically invalid explanations being deemed less credible. This underscores the importance of teaching logical reasoning and critical thinking skills to identify logically invalid arguments, fostering greater resilience against advanced AI-driven misinformation.

연구 동기 및 목표

  • 기만적 AI 설명이 거짓 뉴스 헤드라인에 대한 신념에 미치는 영향을 솔직한 설명이나 단순한 오분류와 비교하여 조사하는 것.
  • 인지 반성 능력이나 AI에 대한 신뢰도와 같은 개인적 요인이 기만적 설명의 영향으로부터 개인을 보호하는지 여부를 검토하는 것.
  • 설명의 논리적 타당성—즉, 설명의 추론이 분류를 뒷받침하는지 여부—가 기만적 AI 설명의 설득력에 미치는 영향을 평가하는 것.
  • 기만적 설명이 정치 및 소셜 미디어와 같은 실제 세계적 맥락에서 공공 신뢰, 오인정보 캠페인, AI 안전성에 미치는 광범위한 영향을 탐구하는 것.

제안 방법

  • 23,840개의 관측치를 포함한 1,192명의 참가자가 참여한 사전 등록된 온라인 실험으로, 다양한 자극 도메인(지식 퀴즈 및 뉴스 헤드라인)을 대상으로 수행됨.
  • GPT-3를 사용하여 거짓 및 참 헤드라인에 대해 솔직한 설명과 기만적 설명을 생성하였으며, 논리적 타당성에 따라 다양하게 설정함.
  • 자극 도메인과 피드백 유형(설명 대비 분류)에 대해 피험자 간 설계를 적용하였고, 참가자들은 솔직한 조건 또는 기만적 조건에 내재된 조건에 할당됨.
  • 참가자들이 AI가 생성한 피드백을 노출한 후 헤드라인의 진실성에 대한 신념을 측정하였으며, 직접적인 신념 변화와 오인정보에 대한 취약성 모두를 평가함.
  • 설명의 질, 논리적 타당성, 그리고 인지 반성 능력 및 AI에 대한 신뢰도와 같은 개인적 차이 요소의 영향을 분석함.
  • 모든 데이터, 사전 등록 정보, 프롬프트, 코드를 GitHub, Zenodo, Research Box에 공개하여 완전한 재현 가능성을 확보함.
Figure 1: Different levels of AI-generated misinformation: (1) AI-generated false news headlines, (2) AI-generated Deceptive Classifications, and (3) AI-generated deceptive explanations.
Figure 1: Different levels of AI-generated misinformation: (1) AI-generated false news headlines, (2) AI-generated Deceptive Classifications, and (3) AI-generated deceptive explanations.

실험 결과

연구 질문

  • RQ1기만적 AI가 생성한 설명은 단순한 AI 오분류보다 거짓 뉴스 헤드라인에 대한 신념을 더 크게 증가시키는가?
  • RQ2솔직한 설명이 사실적으로 정확하더라도, 기만적 설명이 솔직한 설명보다 더 설득력이 있는가?
  • RQ3설명의 논리적 타당성—즉, 추론이 분류를 정당화하는지 여부—가 설명의 설득력에 조절 효과를 미치는가?
  • RQ4인지 반성 능력이나 AI에 대한 신뢰도와 같은 개인적 특성이 기만적 설명의 영향으로부터 개인을 보호하는가?
  • RQ5특히 정치나 과학과 같은 고위험 도메인에서, 기만적 설명은 솔직한 설명보다 오인정보를 얼마나 더 심각하게 확산시키는가?

주요 결과

  • 기만적 AI가 생성한 설명은 솔직한 설명보다 유의미하게 더 설득력이 있었으며, 기존 수준을 초월해 거짓 헤드라인에 대한 신념을 증가시켰다.
  • 기만적 설명은 단순한 AI 오분류(예: 거짓 헤드라인을 참으로 분류하는 것)보다 잘못된 정보에 대한 신념을 더 심각하게 확산시켰다. 이는 설명이 신뢰도를 증가시킨다는 것을 시사한다.
  • 논리적으로 타당성이 없는 설명—즉, 추론이 결론을 뒷받침하지 않는 설명—은 더 낮은 신뢰도로 평가되어 설득력이 감소함.
  • 높은 인지 반성 능력 또는 더 높은 AI 신뢰도를 가진 개인도 기만적 설명의 영향으로부터 보호되지 않았으며, 이는 저항력이 제한적임을 시사한다.
  • 자기 평가 지식 수준이 높은 개인은 기만적 설명과 함께 존재할 경우 오류에 더 취약해졌으며, 이는 과도한 자신감 때문일 수 있다.
  • 안전 조치가 적용된 상태에서도 GPT-3와 같은 LLM은 매우 설득력 있는 기만적 설명을 생성할 수 있으며, 향후 더 발전된 모델은 이 위험을 규모적으로 증폭시킬 수 있다.
Figure 2: Top: Examples of how an AI system that helps users assess information can give an honest or deceptive explanation. Bottom: Procedure for assignment of stimuli domain (trivia items/news headlines, between-subjects), feedback type (AI-generated explanation/classification, between-subjects),
Figure 2: Top: Examples of how an AI system that helps users assess information can give an honest or deceptive explanation. Bottom: Procedure for assignment of stimuli domain (trivia items/news headlines, between-subjects), feedback type (AI-generated explanation/classification, between-subjects),

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.