[论文解读] From Instructions to Intrinsic Human Values -- A Survey of Alignment Goals for Big Models
本文综述了大型语言模型(LLMs)对齐目标的发展,追溯其从人类指令到人类偏好,最终到内在人类价值观的演变过程。文章认为,将LLMs与核心人类价值观(如有益、诚实、无害)对齐是最根本且可持续的目标,并指出了实现稳健、基于价值观对齐的关键挑战与资源。
Big models, exemplified by Large Language Models (LLMs), are models typically pre-trained on massive data and comprised of enormous parameters, which not only obtain significantly improved performance across diverse tasks but also present emergent capabilities absent in smaller models. However, the growing intertwining of big models with everyday human lives poses potential risks and might cause serious social harm. Therefore, many efforts have been made to align LLMs with humans to make them better follow user instructions and satisfy human preferences. Nevertheless, `what to align with' has not been fully discussed, and inappropriate alignment goals might even backfire. In this paper, we conduct a comprehensive survey of different alignment goals in existing work and trace their evolution paths to help identify the most essential goal. Particularly, we investigate related works from two perspectives: the definition of alignment goals and alignment evaluation. Our analysis encompasses three distinct levels of alignment goals and reveals a goal transformation from fundamental abilities to value orientation, indicating the potential of intrinsic human values as the alignment goal for enhanced LLMs. Based on such results, we further discuss the challenges of achieving such intrinsic value alignment and provide a collection of available resources for future research on the alignment of big models.
研究动机与目标
- 识别并分析指导大型语言模型(LLMs)发展的根本对齐目标。
- 研究对齐目标从表面指令到深层人类价值观的演变过程。
- 论证内在人类价值观是LLM对齐最根本且可持续的目标。
- 突出LLMs在实现稳定、有效且全面的价值对齐方面面临的关键挑战。
- 提供一份精选的公开资源列表,以支持未来LLM对齐研究。
提出的方法
- 对现有对齐工作进行全面调查,将对齐目标分为三个层次:人类指令、人类偏好和人类价值观。
- 通过系统性文献回顾,分析这三个层次中对齐目标的定义与评估方法。
- 追溯对齐目标从任务特定指令到通用偏好,最终到内在价值观的演变过程。
- 提出宪法AI(Constitutional AI)和上下文学习(in-context learning)是实现直接价值对齐的有前景路径,尽管目前仍受限于对代理演示的依赖。
- 强调需要开发自动、可扩展且全面的评估基准,以在多个难度层级上评估对齐效果。
- 指出需要开发直接优化价值原则而非间接行为信号的对齐算法。
实验结果
研究问题
- RQ1当前大语言模型研究中主要使用的对齐目标是什么?这些目标随时间如何演变?
- RQ2为何将LLMs与内在人类价值观对齐被认为比与指令或偏好对齐更具根本性和可持续性?
- RQ3在LLMs中实现与内在人类价值观稳定且有效的对齐面临哪些关键挑战?
- RQ4如何在多样化的价值相关场景(包括伦理困境)中实现对齐的全面且自动化的评估?
- RQ5有哪些方法和资源可用于支持未来将LLMs与核心人类价值观对齐的研究?
主要发现
- 对齐目标已从表面指令演变为更广泛的偏好,最终发展为有益、诚实、无害等内在人类价值观。
- 直接与内在价值观对齐比依赖人类反馈或演示等代理信号更具鲁棒性和可持续性。
- 当前的评估基准仍成本高昂且不可扩展,因其严重依赖人工标注,凸显了对自动化、可靠度量标准的迫切需求。
- 基于人类反馈的强化学习(RLHF)和监督微调(SFT)因数据噪声和价值维度覆盖不全,难以实现稳定的对齐。
- 宪法AI通过使用明确定义的价值原则生成训练数据展现出前景,但模型仍通过演示学习而非内化原则。
- 未来的对齐算法必须直接将模型与价值原则对齐,并能适应跨文化与语境下多元且不断演变的价值体系。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。