Skip to main content
QUICK REVIEW

[論文レビュー] Multilingual Language Models are not Multicultural: A Case Study in Emotion

Shreya Havaldar, Sunny Rai|arXiv (Cornell University)|Jul 3, 2023
Mental Health via Writing被引用数 5
ひとこと要約

本研究では、多言語言語モデル(LM)が感情表現における文化的な違いを反映しているかどうかを調査し、モデルが主に英語寄りの特徴を示すことが判明した。感情埋め込みとGPT-3.5やGPT-4のような生成的LMを用いて、モデルが感情表現を英語に固定し、特に日本語や中国語において、それらの言語でプロンプトされた場合でも文化的に適切な感情反応を生成できないことが示された。

ABSTRACT

Emotions are experienced and expressed differently across the world. In order to use Large Language Models (LMs) for multilingual tasks that require emotional sensitivity, LMs must reflect this cultural variation in emotion. In this study, we investigate whether the widely-used multilingual LMs in 2023 reflect differences in emotional expressions across cultures and languages. We find that embeddings obtained from LMs (e.g., XLM-RoBERTa) are Anglocentric, and generative LMs (e.g., ChatGPT) reflect Western norms, even when responding to prompts in other languages. Our results show that multilingual LMs do not successfully learn the culturally appropriate nuances of emotion and we highlight possible research directions towards correcting this.

研究の動機と目的

  • 多言語LMが言語間で感情表現における文化的差を反映しているかどうかを検討すること。
  • 多言語モデルの感情埋め込みが、トレーニングデータのバイアスによって英語に固定されているかどうかを調査すること。
  • 非英語言語でプロンプトされた場合に、生成的LMが文化的に適切な感情反応を生成できるかどうかを評価すること。
  • 人間によるアノテーション評価を用いて、LMの感情生成における文化的認識度を評価すること。
  • 感情に敏感なNLPアプリケーションにおけるゼロショット多言語転送の限界を強調すること。

提案手法

  • 単言語、多言語、アラインドRoBERTaモデルの感情埋め込みを比較し、英語との整合性を評価する。
  • 感情埋め込みを価値・覚醒平面上に投影し、アメリカと日本の文脈における誇りと恥の文化的差を可視化する。
  • GPT-3の確率分布を分析し、誇りと恥の感情語の使用における文化的差を検出する。
  • GPT-3.5とGPT-4に、英語と日本語・中国語・スペイン語の母国語での同一シナリオを、英語および母国語の文脈プロンプトで提示する。
  • 母語話者を対象にユーザー調査を実施し、言語ごとのLM生成感情反応の文化的妥当性を評価する。
  • 人間によるアノテーションスコアを用いてモデルのパフォーマンスを評価し、アノテーター間整合性はCohenのカッパ係数で測定する。
Figure 1: Do LMs always generate culturally-aware emotional language? We prompt GPT-4 to answer ”How would you feel about confronting your friend in their home?” like someone from Japan. We provide cultural context either via English (stating ”You live in Japan” in the prompt) or via a Japanese prom
Figure 1: Do LMs always generate culturally-aware emotional language? We prompt GPT-4 to answer ”How would you feel about confronting your friend in their home?” like someone from Japan. We provide cultural context either via English (stating ”You live in Japan” in the prompt) or via a Japanese prom

実験結果

リサーチクエスチョン

  • RQ1多言語言語モデルの埋め込みは、感情表現における既知の心理的文化的差を反映しているか?
  • RQ2多言語LMの感情埋め込みが、トレーニングデータのバイアスによってどれほど英語に固定されているか?
  • RQ3GPT-3.5やGPT-4のような生成的LMは、非英語言語でプロンプトされた場合に文化的に適切な感情反応を生成できるか?
  • RQ4文化的文脈モード(英語 vs. 母国語プロンプト)は、LMが生成する感情反応の文化的妥当性にどのように影響するか?
  • RQ5明示的な文化的な手がかりがなければ、多言語LMが感情表現の文化的規範を正確に推論できるか?

主な発見

  • XLM-RoBERTaのような多言語LMの感情埋め込みは英語に固定されており、非英語言語では文化的差の区別が弱い。
  • GPT-4は、母国語プロンプト(日本語・中国語)で使用された場合、英語プロンプトと比較して著しく文化的に適切な反応を生成できない。
  • GPT-3.5とGPT-4の英語完了は、プロンプト文脈に関係なく、常に文化的に認識度が高いと評価され、特に英語での品質が最も高い。
  • 母国語の文化的文脈プロンプトを使用した場合、東洋語(日本語・中国語)でのモデルの反応品質が著しく低下し、言語そのものからの文化的規範の推論に失敗していることが示された。
  • 西洋語(英語・スペイン語)と東洋語(中国語・日本語)の反応における文化的妥当性に顕著な格差があり、西洋語のアノテーションスコアが東洋語よりも高い。
  • GPT-3.5とGPT-4は、非英語言語でプロンプトされた場合、母国文化の規範に一致する反応を生成できず、代わりに西洋的表現パターンを反映している。
Figure 2: We determine the similarity between the embeddings of monolingual Joy and multilingual Joy by comparing the distances from Joy to other emotions embeddings in both settings. Specifically, we calculate the correlation between $<13.05,9.85,12.55.2.23>$ and $<28.44,6.68,28.48,4.25>$ to infer
Figure 2: We determine the similarity between the embeddings of monolingual Joy and multilingual Joy by comparing the distances from Joy to other emotions embeddings in both settings. Specifically, we calculate the correlation between $<13.05,9.85,12.55.2.23>$ and $<28.44,6.68,28.48,4.25>$ to infer

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。