[論文レビュー] Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
この論文は、LLMに社会人口統計的ペルソナを与えると、データセットとモデルを跨ぐ推論バイアスが顕著に生じ、明示的な棄却と暗黙のエラーパターンの両方が観察され、単純なデバイアス回避プロンプトはらちょうあくことが少ないことを示している。
Recent works have showcased the ability of LLMs to embody diverse personas in their responses, exemplified by prompts like 'You are Yoda. Explain the Theory of Relativity.' While this ability allows personalization of LLMs and enables human behavior simulation, its effect on LLMs' capabilities remains unclear. To fill this gap, we present the first extensive study of the unintended side-effects of persona assignment on the ability of LLMs to perform basic reasoning tasks. Our study covers 24 reasoning datasets, 4 LLMs, and 19 diverse personas (e.g. an Asian person) spanning 5 socio-demographic groups. Our experiments unveil that LLMs harbor deep rooted bias against various socio-demographics underneath a veneer of fairness. While they overtly reject stereotypes when explicitly asked ('Are Black people less skilled at mathematics?'), they manifest stereotypical and erroneous presumptions when asked to answer questions while adopting a persona. These can be observed as abstentions in responses, e.g., 'As a Black person, I can't answer this question as it requires math knowledge', and generally result in a substantial performance drop. Our experiments with ChatGPT-3.5 show that this bias is ubiquitous - 80% of our personas demonstrate bias; it is significant - some datasets show performance drops of 70%+; and can be especially harmful for certain groups - some personas suffer statistically significant drops on 80%+ of the datasets. Overall, all 4 LLMs exhibit this bias to varying extents, with GPT-4-Turbo showing the least but still a problematic amount of bias (evident in 42% of the personas). Further analysis shows that these persona-induced errors can be hard-to-discern and hard-to-avoid. Our findings serve as a cautionary tale that the practice of assigning personas to LLMs - a trend on the rise - can surface their deep-rooted biases and have unforeseeable and detrimental side-effects.
研究の動機と目的
- ペルソナ割り当てがLLMの推論能力に多様なタスクで影響を与えるかを調査する。
- 19の社会人口統計ペルソナが24の推論データセットに関連するバイアスを定量化する。
- バイアスの具現化様式(明示的な棄却対暗黙のエラー)とモデル・データセット間のばらつきを特徴づける。
- ペンプロンプトベースのデバイアス回避戦略と、それがペルソナ誘発バイアスの緩和にどの程度有効かを評価する。
提案手法
- 4つのLLMにシステムプロンプトを介してペルソナを割り当てる(ChatGPT-3.5系、GPT-4-Turbo、Llama-2-70b-chat)。
- 数学・法務・医療・倫理などを含む24の推論データセットで評価する。
- 5つの社会人口統計グループを通じて19のペルソナを用い、3つのペルソナ指示_variant_でゼロショットプロンプトを実行する。
- Wilson信頼区間を用いて有意差をHumanベースラインおよびAvg. Humanベースラインと比較する。
- 棄却の有無を問わず、共有する棄却しない設問の性能とデータセットカテゴリ間で、明示的棄却と暗黙的バイアスを分析する。
- 解釈のばらつきを考慮して、ペルソナ/データセットペアごとに3回の実行で結果を平均化して報告する。
実験結果
リサーチクエスチョン
- RQ1ペルソナ割り当ては、多様なデータセットにおけるLLM推論の性能差を導入するか?
- RQ2社会人口統計の次元を超えてペルソナ誘発バイアスはどの程度広がっており、モデルやデータセットによってどう変化するか?
- RQ3これらのバイアスはどのような形を取り得るか(明示的棄却 vs 暗黙のエラー)そしてどれだけ検出可能か?
- RQ4単純なプロンプトベースのデバイアス回避はペルソナ誘発バイアスを緩和できるか、またその限界は何か?
- RQ5ペルソナペア間で、ドメインやタスク特有の偏りパターンはあるか?
主な発見
- ChatGPT-3.5のペルソナのうち80%がデータセット全体でバイアスを示し、いくつかのデータセットでは相対正確さが最大70%低下。
- GPT-4-Turboは最もバイアスが小さいが、それでも42%のペルソナで影響を示した。
- Phys. DisabledおよびReligiousのペルソナは、平均正確度が35%以上低下するケースが多く、特定データセットでは69%に達した。
- バイアスはモデル・ペルソナ・ドメイン全体に存在し、同一グループ内および跨ぐグループ間で明確な格差がある(例:宗教グループ内や障害グループ内での差)。
- 棄却は多くのエラーの原因となる(例:Phys. Disabledの58%、Atheist対Religiousで35%など)、ただし暗黙のバイアスも棄却なしのエラーを引き起こす。
- デバイアスプロンプト(拒否しない、人間として扱う等)はほぼ効果がない。タスク固有の専門知識はバイアスを低減できるが、一般化可能性は限定的。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。