Skip to main content
QUICK REVIEW

[논문 리뷰] Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Stephen Casper, Xander Davies|arXiv (Cornell University)|2023. 07. 27.
Software Reliability and Analysis Research인용 수 89
한 줄 요약

논문은 인간 피드백, 보상 모델링, 정책 최적화 전반에 걸친 RLHF의 미해결 문제와 근본적 한계를 고찰하고, 더 넓은 안전성 및 거버넌스 함의를 논의한다.

ABSTRACT

Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of-the-art large language models (LLMs). Despite this popularity, there has been relatively little public work systematizing its flaws. In this paper, we (1) survey open problems and fundamental limitations of RLHF and related methods; (2) overview techniques to understand, improve, and complement RLHF in practice; and (3) propose auditing and disclosure standards to improve societal oversight of RLHF systems. Our work emphasizes the limitations of RLHF and highlights the importance of a multi-faceted approach to the development of safer AI systems.

연구 동기 및 목표

  • RLHF 전반에서 인간 피드백, 보상 모델링, 정책 최적화에 관한 구체적 도전을 분류한다.
  • RLHF를 넘어서는 접근이 필요한 근본적 한계와 해결 가능성이 있는 문제를 구분한다.
  • RLHF를 더 큰 안전 프레임워크와 거버넌스 고려와 통합하는 방안을 논의한다.

제안 방법

  • 피드백 수집, 보상 모델링, 정책 최적화의 세 가지 주요 영역으로 기존의 RLHF 도전 과제를 검토 및 분류한다.
  • 근본적 한계와 해결 가능한 개선점을 요약하고, 부록 B에서 구분을 설명한다.
  • RLHF 중심의 접근 방식이 더 넓은 안전 및 감시 전략에 어떻게 맞춰지는지 종합한다.
Figure 1: (Top) Reinforcement Learning from Human Feedback. Gray, rounded boxes correspond to outputs (e.g., text), and colored diamonds correspond to evaluations. (Bottom) Our taxonomy for challenges with RLHF. We divide challenges with RLHF into three main types: challenges with obtaining quality
Figure 1: (Top) Reinforcement Learning from Human Feedback. Gray, rounded boxes correspond to outputs (e.g., text), and colored diamonds correspond to evaluations. (Bottom) Our taxonomy for challenges with RLHF. We divide challenges with RLHF into three main types: challenges with obtaining quality

실험 결과

연구 질문

  • RQ1RLHF의 주요 도전 카테고리는 피드백, 보상 모델, 정책이며, RLHF 내에서 근본적인지 여부와 해결 가능한지 여부는 무엇인가?
  • RQ2RLHF가 추가 안전 조치 및 거버넌스 관행으로 보완되어 사회적 감독을 어떻게 개선할 수 있는가?
  • RQ3잘못 정렬된 인간, 데이터 품질, 보상 해킹이 RLHF 배치에 어떤 함의를 갖는가?
  • RQ4RLHF 시스템의 투명성과 책임성을 향상시킬 수 있는 감사 및 공시 기준은 무엇인가?

주요 결과

  • RLHF는 피드백 수집, 보상 모델링, 정책 최적화의 단계에서 도전에 직면한다.
  • 많은 한계는 근본적이며 RLHF를 넘어서는 접근이 필요하다.
  • 감독, 데이터 품질, 평가의 복잡성은 오정렬, 편향, 조작과 같은 지속적 위험을 초래한다.
  • RLHF는 단일 정렬 방법으로 의존하기보다는 다층적 안전 프레임워크에 통합되어야 한다.
  • RLHF 기반 시스템의 책임성을 높일 수 있는 거버넌스 및 투명성 고려사항이 있다.
Figure 2: An example of RLHF for finetuning chatbots with binary preference feedback. Humans indicate which example between a pair they prefer. A reward model is trained using each example pair to provide rewards that reflect the human’s decisions. Finally, the LLM policy is finetuned using the rewa
Figure 2: An example of RLHF for finetuning chatbots with binary preference feedback. Humans indicate which example between a pair they prefer. A reward model is trained using each example pair to provide rewards that reflect the human’s decisions. Finally, the LLM policy is finetuned using the rewa

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.