Skip to main content
QUICK REVIEW

[论文解读] Diminished Diversity-of-Thought in a Standard Large Language Model

Peter S. Park, Philipp Schoenegger|arXiv (Cornell University)|Feb 13, 2023
Computational and Text Analysis Methods被引用 10
一句话总结

论文将GPT-3.5作为社会科学重复研究中人类参与者的代理进行测试,并记录一种现象:某些提示会导致答案几乎无变化,这挑战LLMs作为人类主体普遍替代的有效性。

ABSTRACT

We test whether Large Language Models (LLMs) can be used to simulate human participants in social-science studies. To do this, we run replications of 14 studies from the Many Labs 2 replication project with OpenAI's text-davinci-003 model, colloquially known as GPT3.5. Based on our pre-registered analyses, we find that among the eight studies we could analyse, our GPT sample replicated 37.5% of the original results and 37.5% of the Many Labs 2 results. However, we were unable to analyse the remaining six studies due to an unexpected phenomenon we call the "correct answer" effect. Different runs of GPT3.5 answered nuanced questions probing political orientation, economic preference, judgement, and moral philosophy with zero or near-zero variation in responses: with the supposedly "correct answer." In one exploratory follow-up study, we found that a "correct answer" was robust to changing the demographic details that precede the prompt. In another, we found that most but not all "correct answers" were robust to changing the order of answer choices. One of our most striking findings occurred in our replication of the Moral Foundations Theory survey results, where we found GPT3.5 identifying as a political conservative in 99.6% of the cases, and as a liberal in 99.3% of the cases in the reverse-order condition. However, both self-reported 'GPT conservatives' and 'GPT liberals' showed right-leaning moral foundations. Our results cast doubts on the validity of using LLMs as a general replacement for human participants in the social sciences. Our results also raise concerns that a hypothetical AI-led future may be subject to a diminished diversity-of-thought.

研究动机与目标

  • 评估标准LLM(GPT-3.5)是否可以模拟社科重复研究中的人类参与者。
  • 在多项任务上衡量相对于Many Labs 2项目的复制成功。
  • 识别并描述削弱LLM回答多样性的现象。

提出的方法

  • 使用 OpenAI GPT-3.5(text-davinci-003)复制14项 Many Labs 2 研究。
  • 注册分析以比较GPT输出与原始结果及Many Labs 2结果。
  • 分析八项可分析研究并报告复制率;记录意外的近零变异现象。
  • 开展探索性后续研究,改变人口统计细节和提示顺序以测试回答的鲁棒性。
  • 在不同提示条件下检查道德基础理论调查结果。

实验结果

研究问题

  • RQ1GPT-3.5 能否复制原始 Many Labs 2 结果的相当一部分?
  • RQ2GPT-3.5 的回答是否表现出足够的多样性,以被视为对人类参与者的有效代理?
  • RQ3出现哪些现象(如“正确答案”效应)会限制将LLM用于社会科学重复研究?
  • RQ4GPT-3.5 派生结论对人口统计和答案顺序提示的改变有多鲁棒?
  • RQ5关于思维多样性的AI主导未来情景有哪些影响?

主要发现

  • GPT-3.5在八项可分析研究中复制了原始结果的37.5%。
  • GPT-3.5在Many Labs 2结果中复制了37.5%。
  • 一个意外的“正确答案”效应导致对微妙问题的回答出现零变异或近零变异。
  • 在探索性后续研究中,“正确答案”对出现在提示之前的人口统计变量具有鲁棒性。
  • 在道德基础理论复制中,GPT-3.5在逆序案例中被识别为保守的比例为99.6%,在逆序中的自由派比例为99.3%,但两组都表现出倾向右派的道德基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。