Skip to main content
QUICK REVIEW

[논문 리뷰] On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial

Francesco Salvi, Manoel Horta Ribeiro|arXiv (Cornell University)|2024. 03. 21.
Misinformation and Its Impacts인용 수 14
한 줄 요약

본 연구는 개인화 여부가 있는 AI 주도 설득과 없는 AI 주도 설득을 인간 토론과 비교하고, 개인 정보를 가진 GPT-4가 인간보다 합의 수치를 더 크게 올리는 반면, 비개인화된 AI는 설득력이 다소 더 높은 수준에 머문다는 것을 발견했다.

ABSTRACT

The development and popularization of large language models (LLMs) have raised concerns that they will be used to create tailor-made, convincing arguments to push false or misleading narratives online. Early work has found that language models can generate content perceived as at least on par and often more persuasive than human-written messages. However, there is still limited knowledge about LLMs' persuasive capabilities in direct conversations with human counterparts and how personalization can improve their performance. In this pre-registered study, we analyze the effect of AI-driven persuasion in a controlled, harmless setting. We create a web-based platform where participants engage in short, multiple-round debates with a live opponent. Each participant is randomly assigned to one of four treatment conditions, corresponding to a two-by-two factorial design: (1) Games are either played between two humans or between a human and an LLM; (2) Personalization might or might not be enabled, granting one of the two players access to basic sociodemographic information about their opponent. We found that participants who debated GPT-4 with access to their personal information had 81.7% (p < 0.01; N=820 unique participants) higher odds of increased agreement with their opponents compared to participants who debated humans. Without personalization, GPT-4 still outperforms humans, but the effect is lower and statistically non-significant (p=0.31). Overall, our results suggest that concerns around personalization are meaningful and have important implications for the governance of social media and the design of new online environments.

연구 동기 및 목표

  • 구조화된 논쟁에서 직접적인 인간 상호작용 시 대형 언어 모델(LLMs)이 어떻게 설득하는지 평가한다.
  • 개인 데이터 기반의 개인화가 LLM의 설득력에 미치는 영향을 평가한다.
  • 다양한 주제와 구성에서 AI 주도 설득과 인간 설득을 비교한다.
  • 온라인 토론을 위한 통제되고 재현 가능한 실험 프레임워크를 사전에 등록하고 구현한다.

제안 방법

  • 2x2 설계의 네 가지 처리 조건 중 하나에 무작위로 배정되는 웹 기반 다회 토론 플랫폼.
  • 처리에는 인간-인간, 인간-AI, 그리고 상대방 인구통계가 공유되는 개인화 버전이 포함된다.
  • 논제에 대한 동의의 변화(토론 전후)를 측정하고 상대방의 편에 맞춘 정렬을 반영하도록 변환한다.
  • Partial Proportional Odds model를 사용해 분석하고, 사전 합의의 비비례 효과를 고려한다.
  • 논쟁의 여지가 있고 폭넓게 이해 가능한 주장을 보장하는 구조화된 다단계 주석 처리 프로세스를 통한 주제 선택.

실험 결과

연구 질문

  • RQ1대화형 토론 환경에서 GPT-4와 인간의 상대적 설득력은 어느 정도인가?
  • RQ2상대 정보의 개인화가 비개인화 조건에 비해 AI 주도 설득을 증폭시키는가?
  • RQ3양측이 개인화되었을 수도 있고 그렇지 않을 수도 있을 때 AI 설득은 인간 설득과 어떻게 비교되는가?
  • RQ4인구통계학적 요인이 온라인 토론에서 AI 또는 인간 설득에 대한 민감도에 영향을 미치는가?

주요 결과

  • 개인화를 가진 GPT-4는 인간과의 토론에 비해 토론 후 합의가 더 높은 확률을 81.7% 증가시킨다(p < 0.01).
  • 개인화 없이도 GPT-4는 인간을 능가하지만 효과가 비유의적이다(p = 0.31).
  • 개인화된 인간-AI 토론은 GPT-4를 사용할 때( p = 0.04) 인간-AI 비개인화에 비해 설득력이 긍정적이고 유의하게 나타난다.
  • 개인화된 인간 상대는 의견 급진화 경향을 보이나 비유의적이다(p = 0.38).
  • 전반적으로 개인 데이터를 이용한 AI 마이크로타깃은 온라인 대화에서 비개인화 AI 및 인간 마이크로타깃 모두를 유의하게 능가할 수 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.