[論文レビュー] Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
この論文はLLMを裁く judge として用いる12の偏見を定義し、これらの偏見を定量化・分析する自動化フレームワーク Calm を紹介し、6つのLLMを評価して持続する偏見と信頼性の限界を明らかにする。
LLM-as-a-Judge has been widely utilized as an evaluation method in various benchmarks and served as supervised rewards in model training. However, despite their excellence in many domains, potential issues are under-explored, undermining their reliability and the scope of their utility. Therefore, we identify 12 key potential biases and propose a new automated bias quantification framework-CALM-which systematically quantifies and analyzes each type of bias in LLM-as-a-Judge by using automated and principle-guided modification. Our experiments cover multiple popular language models, and the results indicate that while advanced models have achieved commendable overall performance, significant biases persist in certain specific tasks. Empirical results suggest that there remains room for improvement in the reliability of LLM-as-a-Judge. Moreover, we also discuss the explicit and implicit influence of these biases and give some suggestions for the reliable application of LLM-as-a-Judge. Our work highlights the need for stakeholders to address these issues and remind users to exercise caution in LLM-as-a-Judge applications.
研究の動機と目的
- LLMが裁く役割を担う際に影響を与え得る12の偏見を定義し、分類する。
- 自動化された摂動ベースの Calm を提案し、裁判官の偏見を定量化する。
- 複数の LLM を評価して、偏見下での判断の頑健性と信頼性を評価する。
- ベンチマークと報酬における LLM-as-a-Judge の信頼性ある展開のためのガイダンスを提供する。
提案手法
- Calm (Comprehensive Assessment of Language Model Judge Biases) を4つの構成要素(偏見分類法、多様な評価データセット、偏見別の指標、自動化された摂動)で導入する。
- 原理指向の摂動 g(·) によって R または I に偏見を注入し、一貫性を評価する攻撃・検出アプローチを用いる。
- 事実関連、改良志向、整合性データセットを横断するスコアリングと対比較Judgingタスクの両方を適用する。
- 偏見の影響を定量化するために、Robustness Rate (RR)、Consistency Rate (CR)、および正確さの指標などの指標を定義する。
- 6つのLLM(ChatGPT、GPT-4-Turbo、GPT-4o、Claude-3.5、GLM-4、Qwen2)を複数の偏見シナリオ下で評価する。
実験結果
リサーチクエスチョン
- RQ1裁判官として用いられる際に影響を与える12の異なる偏見とは何か?
- RQ2自動化された摂動はLLMベースの判断の頑健性と信頼性をどのように定量化できるのか?
- RQ3最新のLLMは異なる judging タスクやデータセットを横断して持続的な偏見を示すのか?
- RQ4実践における LLM-as-a-Judge の信頼性を向上させるためのガイダンスと緩和戦略は何か?
主な発見
- 偏見は judging の頑健性に大きく影響する;強力なモデルでもタスクとデータセットに依存した脆弱性を示す。
- 整合性データは事実関連データよりも偏見効果が強いことを示し、データセットの質が judge の信頼性に影響する。
- Claude-3.5 は一般的に偏見に対する耐性が高いが、いかなるモデルも全ての偏見タイプに対して普遍的に頑健とは言えない。
- 位置、冗長性、および自己強化の偏見が顕著で、自己強化は生成源と評価源の間に明確な乖離を示す。
- Chain-of-thought (CoT) は一部のモデルにとって評価精度を向上させることがあるが、全てのモデルに普遍的には適用できない。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。