[论文解读] Are Emergent Abilities of Large Language Models a Mirage?
本文认为大语言模型中所谓的涌现能力并非由根本的尺度扩展所驱动,而是由指标选择和数据限制所造成。它提出了一个简单的数学模型和三个互补的测试,表明在线性/连续指标、改进统计或跨视觉任务时,涌现能力可能会消失。
Recent work claims that large language models display emergent abilities, abilities not present in smaller-scale models that are present in larger-scale models. What makes emergent abilities intriguing is two-fold: their sharpness, transitioning seemingly instantaneously from not present to present, and their unpredictability, appearing at seemingly unforeseeable model scales. Here, we present an alternative explanation for emergent abilities: that for a particular task and model family, when analyzing fixed model outputs, emergent abilities appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale. Specifically, nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes in model performance. We present our alternative explanation in a simple mathematical model, then test it in three complementary ways: we (1) make, test and confirm three predictions on the effect of metric choice using the InstructGPT/GPT-3 family on tasks with claimed emergent abilities; (2) make, test and confirm two predictions about metric choices in a meta-analysis of emergent abilities on BIG-Bench; and (3) show to choose metrics to produce never-before-seen seemingly emergent abilities in multiple vision tasks across diverse deep networks. Via all three analyses, we provide evidence that alleged emergent abilities evaporate with different metrics or with better statistics, and may not be a fundamental property of scaling AI models.
研究动机与目标
- 质疑涌现能力是否 intrinsic 于模型规模而非测量的结果。
- 提出一个将每个令牌错误、指标选择与观察到的涌现联系起来的数学模型。
- 使用 InstructGPT/GPT-3 验证预测以评估指标对涌现性算术能力的影响。
- 对 BIG-Bench 结果进行元分析以评估涌现能力对指标的依赖性。
- 通过改变评估指标,在视觉任务中展示诱导的涌现能力。
提出的方法
- 给出一个简单的数学模型,其中每个令牌的交叉熵损失 L_CE(N) 与模型规模 N 成幂律关系,从而在非线性或不连续的指标下产生对任务分数的非线性映射。
- 显示非线性指标(如长序列的准确性)和不连续指标(如多项选择等级)如何在小模型到大模型之间造成尖锐的、看起来涌现的转变。
- 证明连续/线性指标(如令牌编辑距离、Brier 分数)在尺度扩大时会产生平滑、可预测的改进,从而削弱涌现行为。
- 在 InstructGPT/GPT-3 上测试三个预测:指标变化揭示平滑改进;更高分辨率的统计显示非零的小模型在非线性指标上的表现;目标长度对性能有可预测的影响。
- 对 BIG-Bench 的结果进行元分析,观察涌现能力仅在少数非线性/不连续指标下出现,在连续指标下消失。
- 通过设计适当的重建与序贯分类任务评估标准,展示如何在不同架构的视觉模型中诱导涌现性能力。
实验结果
研究问题
- RQ1涌现能力是否取决于用于评估模型性能的指标?
- RQ2非线性或不连续指标是否会在模型规模上产生看似锐利的转变,而在线性/连续指标下消失?
- RQ3更高分辨率的统计数据(更多测试数据)是否揭示在具有所谓涌现能力的任务上小模型的非零表现?
- RQ4通过改变评估指标,是否能在非语言领域(视觉)中诱导涌现能力?
- RQ5BIG-Bench 的涌现性主张对指标选择和模型家族的鲁棒性如何?
主要发现
- 涌现能力主要出现在非线性或不连续指标下,如精确字符串匹配(Exact String Match)和多项选择等级(Multiple Choice Grade)。
- 提高测试数据分辨率揭示小模型的超出随机水平的表现,表明在非线性指标下存在非零能力。
- 改用线性或连续指标(如令牌编辑距离或 Brier 分数)会产生平滑、可预测的改进,从而削弱涌现效应。
- 元分析显示涌现能力在大多数指标中并不普遍;两种指标(Multiple Choice Grade 与 Exact String Match)解释了大部分被声称的涌现能力。
- 通过设计合适的评估标准,可以在不同架构的视觉模型中诱导新的涌现样能力。
- 对 InstructGPT/GPT-3 的实证结果支持在改进统计与指标改变后涌现能力会消失。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。