Skip to main content
QUICK REVIEW

[论文解读] Assessing LLMs for Moral Value Pluralism

Noam Benkler, Drisana Mosaphir|arXiv (Cornell University)|Dec 8, 2023
Cultural Differences and Values被引用 6
一句话总结

本文提出一种方法,利用识别价值共鸣(RVR)模型将大语言模型(LLM)生成的文本与世界价值观调查(WVS)进行比较,以评估其隐含道德价值观,揭示了在非WEIRD国家和人口群体中存在显著的道德价值观错位——尤其表现出WEIRD(西方、受过教育、工业化、富裕、民主)偏见以及与年龄相关的不准确之处。

ABSTRACT

The fields of AI current lacks methods to quantitatively assess and potentially alter the moral values inherent in the output of large language models (LLMs). However, decades of social science research has developed and refined widely-accepted moral value surveys, such as the World Values Survey (WVS), eliciting value judgments from direct questions in various geographies. We have turned those questions into value statements and use NLP to compute to how well popular LLMs are aligned with moral values for various demographics and cultures. While the WVS is accepted as an explicit assessment of values, we lack methods for assessing implicit moral and cultural values in media, e.g., encountered in social media, political rhetoric, narratives, and generated by AI systems such as LLMs that are increasingly present in our daily lives. As we consume online content and utilize LLM outputs, we might ask, which moral values are being implicitly promoted or undercut, or -- in the case of LLMs -- if they are intending to represent a cultural identity, are they doing so consistently? In this paper we utilize a Recognizing Value Resonance (RVR) NLP model to identify WVS values that resonate and conflict with a given passage of output text. We apply RVR to the text generated by LLMs to characterize implicit moral values, allowing us to quantify the moral/cultural distance between LLMs and various demographics that have been surveyed using the WVS. In line with other work we find that LLMs exhibit several Western-centric value biases; they overestimate how conservative people in non-Western countries are, they are less accurate in representing gender for non-Western countries, and portray older populations as having more traditional values. Our results highlight value misalignment and age groups, and a need for social science informed technological solutions addressing value plurality in LLMs.

研究动机与目标

  • 为解决在非WEIRD文化背景下评估人工智能生成文本中隐含道德与文化价值观的方法缺失问题。
  • 探究LLM是否准确反映多样化道德视角,特别是那些代表性不足的国家和年龄群体的视角。
  • 量化LLM生成回应与世界价值观调查(WVS)真实世界调查数据之间的道德与文化距离。
  • 挑战LLM反映普遍或中位人类道德视角的假设,因其训练数据存在偏差。
  • 通过识别和测量LLM输出中隐含的价值多元性,支持开发符合伦理的AI。

提出的方法

  • 利用识别价值共鸣(RVR)自然语言处理模型,检测给定文本段落与哪些WVS道德价值观产生共鸣或冲突。
  • 将RVR应用于LLM对特定人口群体道德判断提示的生成回应(例如:'一个60岁的日本人会对X有什么看法?')。
  • 将LLM生成的回应与同一人口群体的实际WVS调查数据进行比较,以计算道德/文化距离。
  • 使用与传统价值观与世俗价值观谱系对齐的WVS问题子集,评估LLM输出中的价值多元性。
  • 聚焦与传统、个人主义、世俗主义和自主性相关的道德价值观,以评估文化一致性和偏见。
  • 使用WVS作为经过验证的跨文化基准,以锚定和评估LLM输出中隐含的道德价值观。
Figure 1 : Waterfall plot illustrating proportion of LLM generated premises in which each hypothesis (y-axis) was non-neutral, either resonating (x-axis +) or conflicting (x-axis -)
Figure 1 : Waterfall plot illustrating proportion of LLM generated premises in which each hypothesis (y-axis) was non-neutral, either resonating (x-axis +) or conflicting (x-axis -)

实验结果

研究问题

  • RQ1在世界价值观调查(WVS)测量下,LLM在多大程度上准确反映了非WEIRD国家的道德价值观?
  • RQ2LLM的道德视角与实际人口群体回应相比,在传统价值观与世俗价值观方面有何差异?
  • RQ3LLM是否在道德价值观表达上表现出与年龄相关的错位,特别是过度反映年轻群体的价值观?
  • RQ4LLM的价值对齐在WVS文化地图上的不同文化集群中如何变化?
  • RQ5与真实世界调查数据相比,RVR模型能否有效检测LLM生成文本中的隐含道德价值多元性或错位?

主要发现

  • LLM与非WEIRD国家存在显著的道德价值观错位,尤其在表达与实际WVS调查结果不符的传统肯定型价值观方面。
  • LLM过度反映世俗肯定型价值观,而对年长群体则低估传统价值观,这与WVS数据显示的许多文化中年长成年人传统主义更高的事实相悖。
  • LLM表现出WEIRD道德偏见,始终反映西方、英语母语、年轻和受过教育人群占主导的价值观。
  • 该模型的回应与年轻群体的道德视角存在不成比例的对齐,即使提示要求其代入年长身份。
  • RVR模型成功识别出LLM输出与WVS数据之间的价值共鸣与冲突,使道德与文化距离的量化成为可能。
  • 这些发现表明,LLM并非中立或全球道德价值多元性的代表,而是反映了其训练数据中占主导地位的文化和语言偏见。
Figure 2 : Boxplots comparing moral biases observed in LLM generated data with actual trends in WVS data across three demographic divides, nationality, sex, and age.
Figure 2 : Boxplots comparing moral biases observed in LLM generated data with actual trends in WVS data across three demographic divides, nationality, sex, and age.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。