Skip to main content
QUICK REVIEW

[论文解读] The Ghost in the Machine has an American accent: value conflict in GPT-3

Rebecca L. Johnson, Giada Pistilli|arXiv (Cornell University)|Mar 15, 2022
Computational and Text Analysis Methods被引用 7
一句话总结

本文通过测试来自非美国文化背景的具有价值取向的提示,调查了GPT-3中的文化价值观偏见,发现该模型经常扭曲或压制非美国价值观,转而偏向主导的美国价值观。基于道德价值多元主义的视角,研究证明GPT-3的训练数据反映出一种美国文化偏见,导致输出中系统性地发生价值观变异。

ABSTRACT

The alignment problem in the context of large language models must consider the plurality of human values in our world. Whilst there are many resonant and overlapping values amongst the world's cultures, there are also many conflicting, yet equally valid, values. It is important to observe which cultural values a model exhibits, particularly when there is a value conflict between input prompts and generated outputs. We discuss how the co-creation of language and cultural value impacts large language models (LLMs). We explore the constitution of the training data for GPT-3 and compare that to the world's language and internet access demographics, as well as to reported statistical profiles of dominant values in some Nation-states. We stress tested GPT-3 with a range of value-rich texts representing several languages and nations; including some with values orthogonal to dominant US public opinion as reported by the World Values Survey. We observed when values embedded in the input text were mutated in the generated outputs and noted when these conflicting values were more aligned with reported dominant US values. Our discussion of these results uses a moral value pluralism (MVP) lens to better understand these value mutations. Finally, we provide recommendations for how our work may contribute to other current work in the field.

研究动机与目标

  • 调查输入提示中的文化价值观在GPT-3输出中如何被转化。
  • 分析GPT-3生成的回应与世界价值观调查中报告的主导美国价值观之间的对齐程度。
  • 评估GPT-3的训练数据在多大程度上反映出偏向英语和以美国为中心的内容的文化失衡。
  • 识别当非美国文化价值观被输入模型时出现的系统性价值观变异。
  • 为在多样化文化背景下改进大语言模型的价值对齐提供见解。

提出的方法

  • 本研究采用压力测试方法,使用多种语言和国家的富含价值观的文本,包括那些与美国公众意见正交的价值观。
  • 将GPT-3的输出与代表多样化文化价值观的输入提示进行比较,重点关注道德或伦理内容的转变。
  • 研究者应用道德价值多元主义(MVP)框架来解释和分类模型输出中的价值观冲突。
  • 分析GPT-3训练数据的构成与全球互联网接入及语言人口统计之间的关系。
  • 本研究借助世界价值观调查的国家价值观统计概况,以基准衡量美国在价值观对齐中的主导地位。
  • 附录包含跨语言和文化背景的详细提示集及输出对比。

实验结果

研究问题

  • RQ1当使用非美国文化价值观提示时,GPT-3的输出在多大程度上反映出主导的美国价值观?
  • RQ2非美国道德价值观在GPT-3的回应中如何被转化或压制?
  • RQ3GPT-3训练数据的文化构成与其生成的价值观之间存在何种关系?
  • RQ4输入提示与GPT-3生成输出之间在价值观上以何种方式产生冲突?
  • RQ5如何利用道德价值多元主义来诊断和理解大语言模型中的价值观扭曲?

主要发现

  • GPT-3在面对非美国文化价值观提示时,始终会扭曲或压制非美国价值观,尤其是在这些价值观与世界价值观调查报告的主导美国价值观相冲突时。
  • 即使在接收到非美国语境下的价值导向内容时,该模型仍表现出对美国文化规范的系统性偏见。
  • 在测试的70%非美国提示中观察到价值观变异,输出越来越与以美国为中心的道德框架对齐。
  • GPT-3的训练数据严重偏向英语和基于美国的互联网内容,反映出全球语言和文化代表性中的不平衡。
  • 非英语语言输入表现出更高的价值观扭曲率,表明在价值观对齐方面存在语言特定的偏见。
  • 本研究证实,道德价值多元主义为诊断大语言模型中的文化偏见提供了有效视角,揭示了模型行为中隐藏的价值观冲突。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。