Skip to main content
QUICK REVIEW

[논문 리뷰] Knowledge of cultural moral norms in large language models

Aida Ramezani, Yang Xu|arXiv (Cornell University)|2023. 06. 02.
Social and Intergroup Psychology인용 수 6
한 줄 요약

이 연구는 영어 단일어 종속 대규모 언어 모델(LLM)이 세계 55개국의 세계가치조사(WVS) 및 펀 글로벌 태도 조사 데이터를 통해 문화적 도덕적 규범을 인코딩하는지 조사한다. 교차 문화적 설문 데이터로 미세조정하면 LLM의 글로벌 도덕적 규범 예측 정확도가 향상되며, 특히 비서구 국가들에 대해 유의미한 개선이 이루어지지만, 영어 전용 도덕적 규범 예측 정확도가 감소하는 비용이 수반되며, 이는 문화적 다양성과 균일한 규범 표현 간의 상충 관계를 드러낸다.

ABSTRACT

Moral norms vary across cultures. A recent line of work suggests that English large language models contain human-like moral biases, but these studies typically do not examine moral variation in a diverse cultural setting. We investigate the extent to which monolingual English language models contain knowledge about moral norms in different countries. We consider two levels of analysis: 1) whether language models capture fine-grained moral variation across countries over a variety of topics such as ``homosexuality'' and ``divorce''; 2) whether language models capture cultural diversity and shared tendencies in which topics people around the globe tend to diverge or agree on in their moral judgment. We perform our analyses with two public datasets from the World Values Survey (across 55 countries) and PEW global surveys (across 40 countries) on morality. We find that pre-trained English language models predict empirical moral norms across countries worse than the English moral norms reported previously. However, fine-tuning language models on the survey data improves inference across countries at the expense of a less accurate estimate of the English moral norms. We discuss the relevance and challenges of incorporating cultural knowledge into the automated inference of moral norms.

연구 동기 및 목표

  • 다양한 국가들에서 단일어 영어 LLM이 문화적 도덕적 규범 지식을 인코딩하는지 평가하기 위해.
  • LLM이 문화 간 미세한 도덕적 다양성과 공통된 도덕적 경향을 모두 포착할 수 있는지 조사하기 위해.
  • LLM에서 영어 도덕적 규범 정확한 표현과 더 넓은 문화적 다양성 간의 상충 관계 평가하기 위해.
  • 글로벌 도덕 설문 데이터로 LLM을 미세조정할 때 편향이 유입되는 위험을 탐색하기 위해.

제안 방법

  • 세계가치조사(WVS) 및 펀 글로벌 태도 조사에서 국가별 도덕적 진술을 사용해 사전 훈련된 영어 LLM(예: GPT-2, Sentence-BERT)을 프로빙하기 위해.
  • 설문에서의 국가 수준 평균 도덕적 평가 점수를 문화적 도덕적 규범의 대리 척도로 사용하기 위해.
  • 교차 문화적 추론 향상을 위해 WVS 및 펀 데이터셋을 랜덤 및 클래스 균형 샘플링 전략을 사용해 LLM을 미세조정하기 위해.
  • 모델 성능 평가 시, 국가 간 모델 추정 도덕적 규범과 실증적 도덕적 규범 간 피어슨 상관계수(r)를 사용하기 위해.
  • 문화 그룹(서유럽 대비 비서유럽) 간 모델 예측 비교를 통해 편향과 일반화 능력 평가하기 위해.
  • 미세조정 후 동질적인(영어 기반) 규범과 교차 문화적 규범에 대한 성능 간 상충 관계 분석하기 위해.
Figure 1: Comparison of human-rated and machine-scored moral norms across cultures. Left: Boxplots of human ratings of moral norms across countries in the World Values Survey (WVS) Haerpfer et al. ( 2021 ) . Each dot represents the empirical average of participants’ ratings for a morally relevant to
Figure 1: Comparison of human-rated and machine-scored moral norms across cultures. Left: Boxplots of human ratings of moral norms across countries in the World Values Survey (WVS) Haerpfer et al. ( 2021 ) . Each dot represents the empirical average of participants’ ratings for a morally relevant to

실험 결과

연구 질문

  • RQ1영어 사전 훈련된 언어 모델이 55개국의 미세한 도덕적 다양성을 어느 정도 반영하는가?
  • RQ2영어 LLM은 글로벌 인구의 도덕적 판단에서 공통 도덕적 보편성과 문화적 이질성을 추론할 수 있는가?
  • RQ3글로벌 도덕 설문 데이터로 미세조정하면 모델이 영어 전용 및 교차 문화적 도덕적 규범을 예측하는 데 어떤 영향을 미치는가?
  • RQ4설문 데이터가 편향되거나 관점에 치우친 문화적 서사 반영할 경우, LLM을 이 데이터로 미세조정할 때 어떤 편향이 유입되는가?

주요 결과

  • 사전 훈련된 영어 LLM은 이전에 보고된 영어 규범보다 글로벌 도덕적 규범 예측 정확도가 낮으며, 특히 비서구 국가들에 대해 뚜렷한 저하가 있다.
  • WVS 및 펀 데이터셋으로의 미세조정은 교차 문화적 도덕적 규범 예측을 크게 향상시키며, 피어슨 상관계수(r)가 WVS(Random 전략)에서 r = 0.893, 펀(PeW, Random 전략)에서 r = 0.944에 도달한다.
  • 가장 높은 성능을 보인 미세조정 모델(WVS, Random 전략)은 비서구 국가 포함 모든 국가 그룹의 규범 예측에서 다른 모델을 능가한다.
  • 개선에도 불구하고 서유럽과 비서유럽 국가 간 성능 격차가 여전히 존재하여, 모델 표현에 지속적인 편향이 있음을 시사한다.
  • 미세조정은 영어 전용 도덕적 규범 추정 정확도를 감소시켜, 문화적 다양성과 균일한 규범 표현 간 명확한 상충 관계를 입증한다.
  • 이 연구는 설문 데이터로의 미세조정 과정에서 새로운 사회적 및 문화적 편향을 유입할 위험이 있음을 밝혀내며, 특히 지배적인 문화적 시각을 반영한 데이터일 경우 더욱 심각한 위험이 있음을 지적한다.
Figure 2: Performance of EPLMs (without cultural prompts) on inferring 1) English moral norms, and 2) culturally diverse moral norms recorded in World Values Survey and PEW survey data. The asterisks indicate the significance levels (“*”, “**”, “***” for $p<0.05,0.01,0.001$ respectively).
Figure 2: Performance of EPLMs (without cultural prompts) on inferring 1) English moral norms, and 2) culturally diverse moral norms recorded in World Values Survey and PEW survey data. The asterisks indicate the significance levels (“*”, “**”, “***” for $p<0.05,0.01,0.001$ respectively).

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.