[论文解读] On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial
本研究比较了带个性化和不带个性化的 AI 驱动说服力与人类辩论的差异,结果发现带有个人信息的 GPT-4 能显著提升认同度比人类更高,而未个性化的 AI 的说服力则仅适度地更强。
The development and popularization of large language models (LLMs) have raised concerns that they will be used to create tailor-made, convincing arguments to push false or misleading narratives online. Early work has found that language models can generate content perceived as at least on par and often more persuasive than human-written messages. However, there is still limited knowledge about LLMs' persuasive capabilities in direct conversations with human counterparts and how personalization can improve their performance. In this pre-registered study, we analyze the effect of AI-driven persuasion in a controlled, harmless setting. We create a web-based platform where participants engage in short, multiple-round debates with a live opponent. Each participant is randomly assigned to one of four treatment conditions, corresponding to a two-by-two factorial design: (1) Games are either played between two humans or between a human and an LLM; (2) Personalization might or might not be enabled, granting one of the two players access to basic sociodemographic information about their opponent. We found that participants who debated GPT-4 with access to their personal information had 81.7% (p < 0.01; N=820 unique participants) higher odds of increased agreement with their opponents compared to participants who debated humans. Without personalization, GPT-4 still outperforms humans, but the effect is lower and statistically non-significant (p=0.31). Overall, our results suggest that concerns around personalization are meaningful and have important implications for the governance of social media and the design of new online environments.
研究动机与目标
- 评估大语言模型在直接的人际互动中相对于有结构的辩论中对人工对手的说服力。
- 评估基于个人数据的个性化对大语言模型说服力的影响。
- 在多种辩论话题和配置上比较 AI 驱动的说服力与人类说服力。
- 预注册并实现一个受控、可重复的在线辩论实验框架。
提出的方法
- 基于网络的多轮辩论平台,随机分配到 2x2 设计中的四种处理条件之一。
- 处理包括人-人、人-AI,以及对手人口统计信息被共享的个性化变体。
- 结果以辩论前后对命题的一致性变化来衡量,并转化为反映与对手立场的一致性的程度。
- 使用部分比例优势顺序回归模型来分析辩论后有序的一致性,同时考虑先前一致性的非成比例效应。
- 通过结构化、分步标注过程选择话题,确保论点可辩论且易于广泛理解。
实验结果
研究问题
- RQ1在互动辩论环境中,GPT-4 相对于人类的相对说服力如何?
- RQ2对手信息的个性化是否相较非个性化条件增强了 AI 驱动的说服力?
- RQ3当双方都可能被个性化时,AI 的说服力与人类的说服力相比如何?
- RQ4人口统计因素是否会影响在在线辩论中对 AI 或人类说服力的易受影响性?
主要发现
- 带个性化的 GPT-4 将辩论后较高一致性的几率相对于与人类辩论提高 81.7%(p < 0.01)。
- 在未进行个性化时,GPT-4 仍然优于人类,但效应不显著(p = 0.31)。
- 在人-AI 个性化辩论中,使用 GPT-4 时相对于非个性化的人- AI 显示出正向且显著的说服效应(p = 0.04)。
- 对手的人身个性化显示出向观点极化的非显著趋势(p = 0.38)。
- 总体而言,带个人数据的 AI 微定向在在线对话中能显著超越非个性化 AI 和人类微定向。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。