Skip to main content
QUICK REVIEW

[논문 리뷰] Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Jiayi Ye, Yanbo Wang|arXiv (Cornell University)|2024. 10. 03.
Law, Economics, and Judicial Systems인용 수 5
한 줄 요약

이 논문은 LLM을 판단자로 사용하는 12가지 편향을 정의하고, 이러한 편향을 정량화하고 분석하는 자동 프레임워크 Calm을 도입하며, 여섯 개의 LLM을 평가하여 지속적인 편향과 신뢰성의 한계를 드러낸다.

ABSTRACT

LLM-as-a-Judge has been widely utilized as an evaluation method in various benchmarks and served as supervised rewards in model training. However, despite their excellence in many domains, potential issues are under-explored, undermining their reliability and the scope of their utility. Therefore, we identify 12 key potential biases and propose a new automated bias quantification framework-CALM-which systematically quantifies and analyzes each type of bias in LLM-as-a-Judge by using automated and principle-guided modification. Our experiments cover multiple popular language models, and the results indicate that while advanced models have achieved commendable overall performance, significant biases persist in certain specific tasks. Empirical results suggest that there remains room for improvement in the reliability of LLM-as-a-Judge. Moreover, we also discuss the explicit and implicit influence of these biases and give some suggestions for the reliable application of LLM-as-a-Judge. Our work highlights the need for stakeholders to address these issues and remind users to exercise caution in LLM-as-a-Judge applications.

연구 동기 및 목표

  • LLM이 심판으로 작용할 때 영향을 미칠 수 있는 12가지 편향을 정의하고 분류한다.
  • 판단 편향을 정량화하기 위한 자동화된 perturbation 기반 프레임워크 Calm를 제안한다.
  • 여러 LLM을 평가하여 편향 하에서 판단의 강건성과 신뢰성을 검토한다.
  • 벤치마크와 보상에서 LLM-as-a-Judge의 신뢰할 수 있는 배치를 위한 가이드를 제공한다.

제안 방법

  • Calm(Comprehensive Assessment of Language Model Judge Biases)와 네 가지 구성요소: 편향 분류 체계, 다양한 평가 데이터셋, 편향별 지표, 자동 perturbations를 도입한다.
  • 원리-가이드 perturbations g(·)를 통해 R 또는 I에 편향을 주입하고 일관성을 판단하는 공격-탐지 방식이다.
  • 사실 관계 데이터, 정교화 인지 데이터, 정렬 데이터셋에 걸쳐 점수 부여 및 쌍대 비교 판단 작업을 모두 적용한다.
  • 편향 영향력을 정량화하기 위한 지표로 Robustness Rate(RR), Consistency Rate(CR), 정확도 지표 등을 정의한다.
  • 여섯 가지 LLM(ChatGPT, GPT-4-Turbo, GPT-4o, Claude-3.5, GLM-4, Qwen2)을 다양한 편향 시나리오 하에서 평가한다.

실험 결과

연구 질문

  • RQ1판단자로 작용할 수 있는 12가지 구별 가능한 편향은 무엇인가?
  • RQ2자동 perturbations가 LLM 기반 판단의 강건성 및 신뢰성을 어떻게 정량화할 수 있는가?
  • RQ3현대의 LLM은 서로 다른 판단 작업과 데이터셋 전반에서 지속적인 편향을 보이는가?
  • RQ4실무에서 LLM-as-a-Judge의 신뢰성을 높이고 완화 전략을 제시할 수 있는 가이드라인은 무엇인가?

주요 결과

  • 편향은 판단의 강건성에 상당한 영향을 미치며, 강력한 모델일지라도 작업 및 데이터셋 의존적으로 취약점을 보인다.
  • 정렬 데이터는 사실 관련 데이터보다 더 강한 편향 효과를 보이며, 데이터셋 품질이 판단자의 신뢰성에 영향을 준다.
  • Claude-3.5는 일반적으로 편향에 대한 회복력이 더 높지만, 어떤 편향 유형에서도 모든 모델이 보편적으로 고정되지는 않는다.
  • 위치, 길이 및 자기강화 편향이 두드러지며, 자기강화 편향은 생성 소스와 평가 소스 간의 분명한 발산을 보인다.
  • CoT(사고과정) 프롬트는 일부 모델의 평가 정확도를 높일 수 있지만 모든 경우에 보편적이지 않다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.