Skip to main content
QUICK REVIEW

[论文解读] Knowledge of cultural moral norms in large language models

Aida Ramezani, Yang Xu|arXiv (Cornell University)|Jun 2, 2023
Social and Intergroup Psychology被引用 6
一句话总结

本研究探讨了单语英语大语言模型(LLMs)是否通过世界价值观调查(World Values Survey)和皮尤全球态度调查数据,在55个国家中编码文化道德规范。在跨文化调查数据上进行微调可提高LLMs对全球道德规范预测的准确性,尤其在非西方国家中表现更佳,但代价是英语特定道德规范的准确性下降,揭示了文化多样性与同质化规范表征之间的权衡。

ABSTRACT

Moral norms vary across cultures. A recent line of work suggests that English large language models contain human-like moral biases, but these studies typically do not examine moral variation in a diverse cultural setting. We investigate the extent to which monolingual English language models contain knowledge about moral norms in different countries. We consider two levels of analysis: 1) whether language models capture fine-grained moral variation across countries over a variety of topics such as ``homosexuality'' and ``divorce''; 2) whether language models capture cultural diversity and shared tendencies in which topics people around the globe tend to diverge or agree on in their moral judgment. We perform our analyses with two public datasets from the World Values Survey (across 55 countries) and PEW global surveys (across 40 countries) on morality. We find that pre-trained English language models predict empirical moral norms across countries worse than the English moral norms reported previously. However, fine-tuning language models on the survey data improves inference across countries at the expense of a less accurate estimate of the English moral norms. We discuss the relevance and challenges of incorporating cultural knowledge into the automated inference of moral norms.

研究动机与目标

  • 评估单语英语LLMs是否在多样化国家中编码文化道德规范知识。
  • 探究LLMs能否捕捉文化间细微的道德差异与共享的道德倾向。
  • 评估在LLMs中准确表征英语道德规范与更广泛文化多样性之间的权衡。
  • 探索在全局道德调查数据上微调LLMs时引入偏见的风险。

提出的方法

  • 使用来自世界价值观调查(WVS)和皮尤全球态度调查的国家特定道德陈述,探测预训练的英语LLMs(如GPT-2、Sentence-BERT)。
  • 将调查中各国的平均道德评分作为文化道德规范的代理指标。
  • 使用随机采样和类别平衡采样策略,在WVS和Pew数据集上对LLMs进行微调,以提升跨文化推理能力。
  • 通过模型估计与实证道德规范之间在各国的皮尔逊相关系数(r)评估模型性能。
  • 比较模型在不同文化群体(西方 vs. 非西方)中的预测表现,以评估偏见与泛化能力。
  • 分析微调后在同质化(英语)规范与跨文化规范上的性能权衡。
Figure 1: Comparison of human-rated and machine-scored moral norms across cultures. Left: Boxplots of human ratings of moral norms across countries in the World Values Survey (WVS) Haerpfer et al. ( 2021 ) . Each dot represents the empirical average of participants’ ratings for a morally relevant to
Figure 1: Comparison of human-rated and machine-scored moral norms across cultures. Left: Boxplots of human ratings of moral norms across countries in the World Values Survey (WVS) Haerpfer et al. ( 2021 ) . Each dot represents the empirical average of participants’ ratings for a morally relevant to

实验结果

研究问题

  • RQ1英语预训练语言模型在多大程度上反映了55个国家间细微的道德差异?
  • RQ2英语LLMs能否推断出全球人口中共享的道德普遍性与文化差异?
  • RQ3在全局道德调查数据上进行微调如何影响模型预测英语特定规范与跨文化道德规范的能力?
  • RQ4当LLMs在反映可能偏颇或视角主导的文化叙事的调查数据上进行微调时,会引入哪些偏见?

主要发现

  • 预训练的英语LLMs对全球道德规范的预测准确性低于以往报告的英语规范准确性,尤其在非西方国家中表现更差。
  • 在WVS和Pew数据集上进行微调显著提升了跨文化道德规范预测能力,皮尔逊相关系数分别达到 r = 0.893(WVS,随机策略)和 r = 0.944(PEW,随机策略)。
  • 表现最佳的微调模型(WVS,随机策略)在预测所有国家群体的规范方面,包括非西方国家,均优于其他模型。
  • 尽管有所改进,西方与非西方国家之间仍存在性能差距,表明模型表征中持续存在偏见。
  • 微调降低了模型对英语特定道德规范估计的准确性,明确揭示了文化多样性与同质化规范表征之间的权衡。
  • 本研究识别出在调查数据上进行微调可能引入新的社会与文化偏见,尤其当数据反映主导文化视角时。
Figure 2: Performance of EPLMs (without cultural prompts) on inferring 1) English moral norms, and 2) culturally diverse moral norms recorded in World Values Survey and PEW survey data. The asterisks indicate the significance levels (“*”, “**”, “***” for $p<0.05,0.01,0.001$ respectively).
Figure 2: Performance of EPLMs (without cultural prompts) on inferring 1) English moral norms, and 2) culturally diverse moral norms recorded in World Values Survey and PEW survey data. The asterisks indicate the significance levels (“*”, “**”, “***” for $p<0.05,0.01,0.001$ respectively).

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。