Skip to main content
QUICK REVIEW

[论文解读] Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Stephen Casper, Xander Davies|arXiv (Cornell University)|Jul 27, 2023
Software Reliability and Analysis Research被引用 89
一句话总结

本论文综述了 RLHF 在来自人类反馈、奖励建模和策略优化方面的开放问题和基本局限,并讨论了更广泛的安全性与治理影响。

ABSTRACT

Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of-the-art large language models (LLMs). Despite this popularity, there has been relatively little public work systematizing its flaws. In this paper, we (1) survey open problems and fundamental limitations of RLHF and related methods; (2) overview techniques to understand, improve, and complement RLHF in practice; and (3) propose auditing and disclosure standards to improve societal oversight of RLHF systems. Our work emphasizes the limitations of RLHF and highlights the importance of a multi-faceted approach to the development of safer AI systems.

研究动机与目标

  • 在 RLHF 的人类反馈、奖励建模和策略优化方面对具体挑战进行分类。
  • 区分可处理的问题与需要超越 RLHF 的基本局限性。
  • 讨论将 RLHF 整合到更广泛的安全框架与治理考量中。

提出的方法

  • 将现有的 RLHF 挑战归纳并分为三个主要领域:反馈收集、奖励建模和策略优化。
  • 总结基本局限性和可改进的可行性,附在附录 B 中作出区分说明。
  • 综合 RLHF 为中心的方法如何融入更广泛的安全与监督策略。
Figure 1: (Top) Reinforcement Learning from Human Feedback. Gray, rounded boxes correspond to outputs (e.g., text), and colored diamonds correspond to evaluations. (Bottom) Our taxonomy for challenges with RLHF. We divide challenges with RLHF into three main types: challenges with obtaining quality
Figure 1: (Top) Reinforcement Learning from Human Feedback. Gray, rounded boxes correspond to outputs (e.g., text), and colored diamonds correspond to evaluations. (Bottom) Our taxonomy for challenges with RLHF. We divide challenges with RLHF into three main types: challenges with obtaining quality

实验结果

研究问题

  • RQ1RLHF 的主要挑战类别有哪些(反馈、奖励模型、策略)?在 RLHF 内哪些是基本的、可实现的可控问题?
  • RQ2RLHF 如何通过额外的安全措施和治理实践来提升社会监督?
  • RQ3人类偏差、数据质量和奖励劫持对 RLHF 部署的影响是什么?
  • RQ4哪些审计与披露标准可以提升 RLHF 系统的透明度与问责性?

主要发现

  • RLHF 在反馈收集、奖励建模和策略优化阶段面临挑战。
  • 许多局限是基本的,需要超越 RLHF 的方法来克服。
  • 监督、数据质量和评估复杂性带来持续风险,如错位、偏见和操控。
  • RLHF 应嵌入到多层安全框架中,而非仅将其作为单一对齐方法。
  • 存在治理与透明度方面的考虑,可提升对基于 RLHF 的系统的问责。
Figure 2: An example of RLHF for finetuning chatbots with binary preference feedback. Humans indicate which example between a pair they prefer. A reward model is trained using each example pair to provide rewards that reflect the human’s decisions. Finally, the LLM policy is finetuned using the rewa
Figure 2: An example of RLHF for finetuning chatbots with binary preference feedback. Humans indicate which example between a pair they prefer. A reward model is trained using each example pair to provide rewards that reflect the human’s decisions. Finally, the LLM policy is finetuned using the rewa

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。