Skip to main content
QUICK REVIEW

[논문 리뷰] Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Jaehun Jung, Faeze Brahman|arXiv (Cornell University)|2024. 07. 25.
Corporate Law and Human RightsBusiness, Management and Accounting인용 수 3
한 줄 요약

이 논문은 신뢰도 추정 기반으로 동적으로 심판 모델을 선택함으로써 LLM 기반 쌍별 평가에서 인간의 일致성을 보장하는 Cascaded Selective Evaluation를 제안한다. 보정된 신뢰도 스코어를 위한 시뮬레이티드 앤오테이터를 사용하여, 79.1%의 커버리지로 80% 이상의 인간 일치도를 달성하면서 GPT-4를 초월하고, 비용을 최대 87.4%까지 절감한다.

ABSTRACT

We present a principled approach to provide LLM-based evaluation with a rigorous guarantee of human agreement. We first propose that a reliable evaluation method should not uncritically rely on model preferences for pairwise evaluation, but rather assess the confidence of judge models and selectively decide when to trust its judgement. We then show that under this selective evaluation framework, human agreement can be provably guaranteed -- such that the model evaluation aligns with that of humans to a user-specified agreement level. As part of our framework, we also introduce Simulated Annotators, a novel confidence estimation method that significantly improves judge calibration and thus enables high coverage of evaluated instances. Finally, we propose Cascaded Selective Evaluation, where we use cheaper models as initial judges and escalate to stronger models only when necessary -- again, while still providing a provable guarantee of human agreement. Experimental results show that Cascaded Selective Evaluation guarantees strong alignment with humans, far beyond what LLM judges could achieve without selective evaluation. For example, on a subset of Chatbot Arena where GPT-4 almost never achieves 80% human agreement, our method, even while employing substantially cost-effective models such as Mistral-7B, guarantees over 80% human agreement with almost 80% test coverage.

연구 동기 및 목표

  • 모델 선호도에 대한 비판적 검토 없이 신뢰도 평가 없이 의존하는 LLM 기반 평가의 증명 가능한 신뢰성 부족 문제를 해결하기 위해.
  • 사용자가 지정한 위험 수준 α에서 인간 일치도를 보장하는 프레임워크를 개발하여, 모델 판단이 인간 공감대에 부합하도록 보장하기 위해.
  • 외부 감시 없이도 LLM 심판의 신뢰도 추정을 향상시켜 평가 커버리지를 극대화하면서도 신뢰성을 훼손하지 않기 위해.
  • 낮은 신뢰도일 경우에만 고급 모델으로 전환하는 계층적 평가를 통해 비용을 절감하면서도 높은 인간 판단 일치도를 유지하기 위해.
  • 선택적 평가와 적절한 신뢰도 보정을 통해 더 약한 모델이 GPT-4와 같은 강력한 모델보다도 더 나은 성능을 낼 수 있음을 입증하기 위해.

제안 방법

  • 신뢰도 추정이 보정된 상태에서 모델의 선호도에 대한 신뢰도가 충분히 높을 경우에만 LLM 심판을 신뢰하는 선택적 평가 프레임워크를 제안한다.
  • 시뮬레이티드 앤오테이터를 도입한다—이것은 문맥 학습을 통해 다양한 앤오테이터 선호도를 시뮬레이션하고, 시뮬레이션 간의 일치 비율을 통해 신뢰도를 계산하는 새로운 비지도적 신뢰도 추정 방법이다.
  • 고정된 시퀀스 테스팅(Bauer, 1991)을 소규모 보정 세트에 적용하여 실질적인 불일치 위험 상한선을 유도함으로써, 인간 일치도를 확률 ≥1−α로 보장한다.
  • 저비용 모델(예: Mistral-7B)이 초깃심판자로 작동하고, 신뢰도가 낮을 경우에만 더 강력한 모델으로 전환하는 캐스케이드드 세레크티브 이밸류에이션을 설계한다. 이 과정에서 인간 일치도 보장을 유지한다.
  • 사용자가 지정한 위험 수준 α에 기반해 최적의 기피 정책을 자동으로 결정함으로써 히우리스틱 기반 모델 선택의 필요성을 제거한다.
  • 모델에 종속되지 않는 접근 방식을 사용하여, 어떤 LLM이라도 신뢰도가 적절히 추정되고 보정된다면 심판자로 활용 가능하다.
Figure 1: Illustration of Cascaded Selective Evaluation. We start with a small, cost-effective model as initial judge, estimate its confidence, and escalate to a stronger model only when the previous judge is not confident. By calibrating when to trust which judge model, our method provides a rigoro
Figure 1: Illustration of Cascaded Selective Evaluation. We start with a small, cost-effective model as initial judge, estimate its confidence, and escalate to a stronger model only when the previous judge is not confident. By calibrating when to trust which judge model, our method provides a rigoro

실험 결과

연구 질문

  • RQ1약한 또는 저비용 모델을 사용하더라도 사용자가 지정한 신뢰 수준에서 LLM 기반 평가가 인간 판단과 일치함을 보장할 수 있는가?
  • RQ2외부 감시나 레이블링 데이터에 의존하지 않고 LLM 심판의 신뢰도 추정을 어떻게 향상시킬 수 있는가?
  • RQ3계층적 평가 프레임워크는 유일하게 최강의 모델을 사용하는 것과 비교해 추론 비용을 절감하면서도 인간 일치도를 유지하거나 향상시킬 수 있는가?
  • RQ4보정된 신뢰도 기반 기피 정책은 길이나 겹침 정도와 같은 얕은 히وري스틱이 아니라 인간의 주관성 인식과 얼마나 잘 일치하는가?
  • RQ5기존의 신뢰도 추정 방법(예: 예측 확률)에 비해 시뮬레이티드 앤오테이터는 보정 수준과 실패 예측에서 더 뛰어나지 않는가?

주요 결과

  • 캐스케이드드 세레크티브 이밸류에이션은 Chatbot Arena 벤치마크에서 심지어 Mistral-7B와 GPT-3.5만을 심판으로 사용하더라도 인간 일치도 80% 이상을 보장한다. 이는 기피 없이 GPT-4를 사용하는 것보다도 뛰어난 성능이다.
  • 이 방법은 Chatbot Arena에서 79.1%의 테스트 커버리지를 달성하면서도 80% 이상의 인간 일치도를 유지한다. 이는 GPT-4가 기피 없이 사용했을 때의 77.8% 일치도를 크게 뛰어넘는 성과이다.
  • 약한 계층(예: Mistral-7B, Mixtral-8×7B, GPT-3.5)을 사용할 경우, GPT-4 대비 평가 비용을 87.4%까지 절감하면서도 인간 일치도 80.3%를 달성한다.
  • 시뮬레이티드 앤오테이터는 예측 확률과 같은 기준 방법에 비해 신뢰도 보정과 실패 예측 능력에서 뚜렷한 향상을 보였다. 특히 예측 확률 방법은 인간 일치도를 과도하게 추정하는 경향이 있다.
  • 기피 정책은 길이 비율이나 토큰 겹침 정도와 같은 얕은 히وري스틱이 아니라 인간의 주관성 인식과 매우 밀접하게 일치한다.
  • 보정 세트에 대한 고정 시퀀스 테스팅을 통해 인간 일치도에 대한 증명 가능한 보장을 제공한다. 이는 임의의 미리보지 않은 인스턴스에 대해서도 불일치 확률이 α 이하로 제한됨을 보장한다.
Figure 2: Reliability plot for confidence estimation methods, using GPT-4 as judge on AlpacaEval. Dashed lines denote perfect calibration, and darker bars denote more samples in the corresponding bins. Simulated Annotators reduces expected calibration error by 50% compared to the baselines, mitigati
Figure 2: Reliability plot for confidence estimation methods, using GPT-4 as judge on AlpacaEval. Dashed lines denote perfect calibration, and darker bars denote more samples in the corresponding bins. Simulated Annotators reduces expected calibration error by 50% compared to the baselines, mitigati

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.