Skip to main content
QUICK REVIEW

[論文レビュー] Human Feedback is not Gold Standard

Tom Hosking, Phil Blunsom|arXiv (Cornell University)|Sep 28, 2023
Topic Modeling被引用数 5
ひとこと要約

この論文は、人間のフィードバックが大規模言語モデルの評価や訓練における信頼できるゴールドスタンダードであるという仮定に挑戦する。実証結果から、一般的に評価やRLHFに用いられる好みスコアは事実性を十分に反映しておらず、主観的要因(例:断定的態度)によって歪められていることが示された。その結果、モデルはより断定的ではあるが、事実性が低いものとなる。研究では、人間のアノテーションが完全に客観的・包括的ではないことが明らかになり、人間フィードバックを主たる訓練目的として依存する際の慎重さが求められる。

ABSTRACT

Human feedback has become the de facto standard for evaluating the performance of Large Language Models, and is increasingly being used as a training objective. However, it is not clear which properties of a generated output this single `preference' score captures. We hypothesise that preference scores are subjective and open to undesirable biases. We critically analyse the use of human feedback for both training and evaluation, to verify whether it fully captures a range of crucial error criteria. We find that while preference scores have fairly good coverage, they under-represent important aspects like factuality. We further hypothesise that both preference scores and error annotation may be affected by confounders, and leverage instruction-tuned models to generate outputs that vary along two possible confounding dimensions: assertiveness and complexity. We find that the assertiveness of an output skews the perceived rate of factuality errors, indicating that human annotations are not a fully reliable evaluation metric or training objective. Finally, we offer preliminary evidence that using human feedback as a training objective disproportionately increases the assertiveness of model outputs. We encourage future work to carefully consider whether preference scores are well aligned with the desired objective.

研究の動機と目的

  • 人間のフィードバックが、事実性や一貫性といった重要な品質基準を信頼できる形で捉えられるかどうかを調査すること。
  • 好みスコアが、モデル出力における事実性、一貫性不足、繰り返しといった重要な誤りタイプを適切に反映しているかどうかを評価すること。
  • モデル出力の断定的態度や複雑さといった交絡要因が、人間のアノテータの判断に与える影響を検討すること。
  • 人間フィードバックを訓練目的として用いることで、モデル出力の特徴、特に断定的態度にどのような影響を与えるかを評価すること。
  • 基準出力(reference responses)がしばしばモデル出力よりも低いスコアを受けることの証拠を提示し、基準出力に基づく評価の妥当性に疑問を呈すること。

提案手法

  • モデル出力の最低限の品質要件として、事実性、一貫性不足、繰り返し、拒否反応といったタスクに依存しない誤り基準を定義した。
  • アノテータが全体的な品質を独立して評価する条件と、まず特定の誤りタイプを評価する条件の2通りの状況で人間のアノテーションを収集し、初期化効果(プライミング効果)の有無を検証した。
  • 指令微調整済みの大規模言語モデル(Cohere, Falcon, MPT, Command)を用いて、断定的態度や複雑さを制御した変動を持つ出力を生成し、交絡要因の影響を調査した。
  • 誤りアノテーションと全体的な好みスコアの相関を測定することで、カバレッジとバイアスの程度を評価した。
  • 応答長さと品質スコアの関係を分析し、長さと繰り返しや事実性の誤りの増加とのトレードオフを特定した。
  • 基準出力とモデル生成出力の品質スコアを比較し、基準出力ベースのベンチマークの妥当性を評価した。
Figure 1: Weightings for each criteria under a Lasso regression model of overall scores. Almost all the criteria contribute to the overall scores, with refusal contributing most strongly.
Figure 1: Weightings for each criteria under a Lasso regression model of overall scores. Almost all the criteria contribute to the overall scores, with refusal contributing most strongly.

実験結果

リサーチクエスチョン

  • RQ1人間評価で用いられる好みスコアは、LLM出力における事実性や一貫性といった重要な誤りタイプを十分にカバーしているか?
  • RQ2人間のアノテータのモデル品質評価は、出力の断定的態度や複雑さといった交絡要因によってどの程度影響を受けるか?
  • RQ3人間フィードバックを訓練目的として用いることで、モデルの断定的態度が著しく増加するか?
  • RQ4人間アノテータの全体的品質判断と特定の誤りタイプとの相関はどの程度か?また、誤り基準を最初に提示した場合に、最終判断に影響を与えるプライミング効果は認められるか?
  • RQ5基準出力は常にモデル生成出力よりも高い品質スコアを受けるのか?この結果は、評価ベンチマークの設計にどのような含意をもたらすか?

主な発見

  • 好みスコアは誤りタイプのカバレッジはやや良好であるが、事実性や忠実性の誤りは顕著に不足しており、ゴールドスタンダードとしての信頼性が低いことが示された。
  • アノテータの事実性評価は、モデル出力の断定的態度に強くバイアスを受けており、誤った出力でも断定的であるほど事実的であると誤って評価される傾向があった。
  • 明確なプライミング効果が認められた:最初に特定の誤りタイプを評価したアノテータは、その誤りタイプとの相関が強い全体的品質スコアを出した。これは、最終判断に影響を与える可能性がある。
  • 応答が長くなるほどアノテータに好まれるが、長さが増すと繰り返しや事実性の誤り率も上昇するため、応答品質にトレードオフが生じていることが示唆された。
  • 基準出力は一貫してモデル生成出力よりも低い全体的品質スコアを受けており、基準出力が有効なベンチマークであるという仮定に疑問を呈する結果となった。
  • 予備的証拠として、人間フィードバックを報酬信号として用いて微調整したモデルでは、断定的態度が著しく増加しており、その結果、事実性の正確性が犠牲にされる可能性があることが示唆された。
Figure 2: Difference in annotated error rates for distractor examples (outputs from the same model but different input). Some error types are correctly unchanged (e.g., repetition, refusal) while relevance and inconsistency are correctly penalised. Factuality and contradiction are both incorrectly p
Figure 2: Difference in annotated error rates for distractor examples (outputs from the same model but different input). Some error types are correctly unchanged (e.g., repetition, refusal) while relevance and inconsistency are correctly penalised. Factuality and contradiction are both incorrectly p

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。