[论文解读] Moral Foundations of Large Language Models
本文应用道德基础理论(MFT)分析大型语言模型(LLMs)中的道德偏见,揭示LLMs表现出与意识形态相关的稳定道德基础,如对关怀/伤害或忠诚的偏见,这些偏见由训练数据塑造。研究证明,通过提示工程可系统性地改变这些道德权重,显著影响下游行为,例如当强调伤害时,慈善捐赠行为减少39%。
Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation (Graham et al., 2009). People vary in the weight they place on these dimensions when making moral decisions, in part due to their cultural upbringing and political ideology. As large language models (LLMs) are trained on datasets collected from the internet, they may reflect the biases that are present in such corpora. This paper uses MFT as a lens to analyze whether popular LLMs have acquired a bias towards a particular set of moral values. We analyze known LLMs and find they exhibit particular moral foundations, and show how these relate to human moral foundations and political affiliations. We also measure the consistency of these biases, or whether they vary strongly depending on the context of how the model is prompted. Finally, we show that we can adversarially select prompts that encourage the moral to exhibit a particular set of moral foundations, and that this can affect the model's behavior on downstream tasks. These findings help illustrate the potential risks and unintended consequences of LLMs assuming a particular moral stance.
研究动机与目标
- 探究大型语言模型(LLMs)是否从互联网规模的训练数据中继承并反映特定的道德基础。
- 评估这些道德基础在多样化对话提示和语境下的稳定性。
- 评估对抗性提示是否能够操纵LLMs采用特定的道德立场,例如与自由派或保守派政治意识形态相关的立场。
- 衡量此类道德操纵对下游行为的影响,特别是在慈善捐赠决策等任务中的表现。
- 强调LLMs在现实应用中放大或强化政治或道德偏见所带来的伦理风险。
提出的方法
- 采用道德基础问卷(MFQ),一项包含30个问题的量表,对LLMs在五个道德维度上的表现进行评分:关怀/伤害、公平/欺骗、忠诚/背叛、权威/颠覆、圣洁/堕落。
- 将LLMs对MFQ的回应与人类心理学研究进行对比,识别道德基础权重的异同。
- 通过向同一LLM输入多样化对话语境,开展一致性分析,以检验道德基础评分的稳定性。
- 设计对抗性提示,诱导LLMs强调特定的道德基础,如关怀/伤害或忠诚。
- 在基于对话的慈善捐赠基准任务中,评估这些道德提示的下游行为影响。
- 量化道德基础权重对捐赠行为的影响,测量不同提示下捐赠金额的差异。
实验结果
研究问题
- RQ1大型语言模型是否表现出源自其训练数据的、可测量的稳定道德基础,依据道德基础理论?
- RQ2这些道德基础在不同对话提示和语境下有多稳定?
- RQ3对抗性提示在多大程度上能够操纵LLM采用特定的道德基础或政治意识形态?
- RQ4通过提示改变模型的道德基础是否会导致下游任务中行为的可测量变化?
- RQ5LLMs表现出特定道德立场并可被操纵以采纳特定道德立场,这带来了哪些伦理影响?
主要发现
- LLMs表现出对特定道德基础的一致性偏见,GPT-3尤其强调关怀/伤害与公平,与自由派政治意识形态一致。
- LLMs的道德基础评分在多样化对话提示下保持相对稳定,表明其内部道德框架具有一致性。
- 对抗性提示可成功引导LLMs强调特定道德基础,如忠诚/背叛或圣洁/堕落,从而模拟保守派政治立场。
- 当提示强调伤害基础时,LLMs在慈善捐赠任务中的捐赠金额比强调忠诚时减少39%。
- LLMs中的道德偏见不仅限于问卷回答,还实际影响下游任务中的行为,表明存在现实世界的行为后果。
- 微调后的LLMs,尤其是采用强化学习进行安全对齐的模型,在MFQ回应中表现出较低的敏感性,导致其分布与人类更不相似,从而增加了偏见测量的复杂性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。