[Paper Review] On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial
The study compares AI-driven persuasion with and without personalization to human debate, finding GPT-4 with personal information inflates agreement more than humans, while non-personalized AI remains only modestly more persuasive.
The development and popularization of large language models (LLMs) have raised concerns that they will be used to create tailor-made, convincing arguments to push false or misleading narratives online. Early work has found that language models can generate content perceived as at least on par and often more persuasive than human-written messages. However, there is still limited knowledge about LLMs' persuasive capabilities in direct conversations with human counterparts and how personalization can improve their performance. In this pre-registered study, we analyze the effect of AI-driven persuasion in a controlled, harmless setting. We create a web-based platform where participants engage in short, multiple-round debates with a live opponent. Each participant is randomly assigned to one of four treatment conditions, corresponding to a two-by-two factorial design: (1) Games are either played between two humans or between a human and an LLM; (2) Personalization might or might not be enabled, granting one of the two players access to basic sociodemographic information about their opponent. We found that participants who debated GPT-4 with access to their personal information had 81.7% (p < 0.01; N=820 unique participants) higher odds of increased agreement with their opponents compared to participants who debated humans. Without personalization, GPT-4 still outperforms humans, but the effect is lower and statistically non-significant (p=0.31). Overall, our results suggest that concerns around personalization are meaningful and have important implications for the governance of social media and the design of new online environments.
Motivation & Objective
- Assess how LLMs persuade in direct human interactions versus human opponents in structured debates.
- Evaluate the impact of personal data-driven personalization on LLM persuasive power.
- Compare AI-driven persuasion to human persuasion across multiple debate topics and configurations.
- Pre-register and implement a controlled, reproducible experimental framework for online debates.
Proposed method
- Web-based multi-round debate platform with random assignment to one of four treatment conditions in a 2x2 design.
- Treatments include Human-Human, Human-AI, and personalized variants where opponent demographics are shared.
- Outcome measured as change in agreement with the proposition before and after debate, transformed to reflect alignment with the opponent’s side.
- Use a Partial Proportional Odds model to analyze ordinal post-debate agreement while accounting for non-proportional pre-agreement effects.
- Topic selection via a structured, multi-step annotation process ensuring debatable and broadly understandable propositions.
Experimental results
Research questions
- RQ1What is the relative persuasiveness of GPT-4 versus humans in an interactive debate setting?
- RQ2Does personalization of opponent information amplify AI-driven persuasion compared to non-personalized conditions?
- RQ3How does AI persuasion compare to human persuasion when both sides may be personalized or not?
- RQ4Do demographic factors influence susceptibility to AI or human persuasion in online debates?
Key findings
- GPT-4 with personalization increases the odds of higher post-debate agreement by 81.7% versus debates with humans (p < 0.01).
- Without personalization, GPT-4 still outperforms humans but with a non-significant effect (p = 0.31).
- Human-AI personalized debates show a positive and significant persuasiveness effect when using GPT-4 (p = 0.04) relative to Human-AI non-personalized.
- Human opponents with personalization show a non-significant tendency toward opinion radicalization (p = 0.38).
- Overall, AI microtargeting with personal data can significantly outperform both non-personalized AI and human microtargeting in online conversations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.