[论文解读] Predictability and Surprise in Large Generative Models
本文认为大型生成模型在一般能力的规模化上表现出随规模平滑且可预测的扩展,而特定能力和输出则突然出现,输入/输出保持开放式,带来部署风险并为政策干预提供指引。
Large-scale pre-training has recently emerged as a technique for creating capable, general purpose, generative models such as GPT-3, Megatron-Turing NLG, Gopher, and many others. In this paper, we highlight a counterintuitive property of such models and discuss the policy implications of this property. Namely, these generative models have an unusual combination of predictable loss on a broad training distribution (as embodied in their "scaling laws"), and unpredictable specific capabilities, inputs, and outputs. We believe that the high-level predictability and appearance of useful capabilities drives rapid development of such models, while the unpredictable qualities make it difficult to anticipate the consequences of model deployment. We go through examples of how this combination can lead to socially harmful behavior with examples from the literature and real world observations, and we also perform two novel experiments to illustrate our point about harms from unpredictability. Furthermore, we analyze how these conflicting properties combine to give model developers various motivations for deploying these models, and challenges that can hinder deployment. We conclude with a list of possible interventions the AI community may take to increase the chance of these models having a beneficial impact. We intend this paper to be useful to policymakers who want to understand and regulate AI systems, technologists who care about the potential policy impact of their work, and academics who want to analyze, critique, and potentially develop large generative models.
研究动机与目标
- 解释大型生成模型的四个显著特征(平滑的一般能力扩展、突然的特定能力扩展、开放式输入、开放式输出)。
- 分析缩放定律如何影响发展激励、部署动机及相关安全挑战。
- 通过新颖实验和现实世界启发的示例说明不可预测性可能带来的危害。
- 讨论政策干预和治理考量,以引导模型发展走向有利的结果。
提出的方法
- 回顾并综合缩放定律文献,展示模型规模、数据、计算和损失之间的幂律关系。
- 给出关于大型语言模型的原创实验,以说明不可预测性带来的危害(例如再犯倾向的提示等)。
- 提供突然的能力出现与开放式行为的定性与定量示例。
- 分析跨不同模型规模的开放式输出和毒性趋势。
- 提出政策干预建议,并讨论产业-学术协同与部署壁垒。
实验结果
研究问题
- RQ1一般能力的扩展是否会随规模、数据和计算呈现可预测的规律?
- RQ2在某些规模下,特定能力是否会突然出现?在何种条件下?
- RQ3开放式输入和输出如何影响对大型模型所带来的危害的预期与缓解?
- RQ4哪些政策与组织干预可以引导大型模型的发展走向更有利的结果?
主要发现
- 缩放定律预测,模型损失会随着更大规模、更多数据和更长的训练而下降,符合幂律关系。
- 特定能力可能在规模达到一定程度时突然出现,呈现出仅从一般规模化看不到的曲棍棒式(Hockey-stick)增益。
- 开放式输入/领域意味着未知能力可能只有在被提示时才显现,增加危害的不可预测性。
- 开放式输出,包括随着模型规模增加而上升的毒性,体现出与能力同样扩张的社会相关风险。
- 在如再犯预测等敏感任务上,大型模型的偏见和危害与现有风险工具相似或更高。
- 存在推动部署的经济、科学和声望动机,以及成本、安全性和缺乏部署标准等障碍。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。