Skip to main content
QUICK REVIEW

[論文レビュー] Are Emergent Abilities of Large Language Models a Mirage?

Rylan Schaeffer, Brando Miranda|arXiv (Cornell University)|Apr 28, 2023
Topic Modeling被引用数 132
ひとこと要約

本論文は、いわゆる emergent abilities が大規模言語モデルの基本的なスケーリングではなく、指標の選択とデータの制約から生じると主張する。単純な数学モデルと三つの補完的なテストを提示し、線形/連続的な指標、より良い統計、あるいは視覚タスクにまたがる場合に emergent abilities が消える可能性を示す。

ABSTRACT

Recent work claims that large language models display emergent abilities, abilities not present in smaller-scale models that are present in larger-scale models. What makes emergent abilities intriguing is two-fold: their sharpness, transitioning seemingly instantaneously from not present to present, and their unpredictability, appearing at seemingly unforeseeable model scales. Here, we present an alternative explanation for emergent abilities: that for a particular task and model family, when analyzing fixed model outputs, emergent abilities appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale. Specifically, nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes in model performance. We present our alternative explanation in a simple mathematical model, then test it in three complementary ways: we (1) make, test and confirm three predictions on the effect of metric choice using the InstructGPT/GPT-3 family on tasks with claimed emergent abilities; (2) make, test and confirm two predictions about metric choices in a meta-analysis of emergent abilities on BIG-Bench; and (3) show to choose metrics to produce never-before-seen seemingly emergent abilities in multiple vision tasks across diverse deep networks. Via all three analyses, we provide evidence that alleged emergent abilities evaporate with different metrics or with better statistics, and may not be a fundamental property of scaling AI models.

研究の動機と目的

  • emergent abilities はモデルのスケーリング固有の性質なのか、それとも評価の測定誤差の産物なのかを検討する。
  • per-token の誤差、指標選択、観測される emergent を結ぶ数学モデルを提案する。
  • InstructGPT/GPT-3 を用いて指標の影響を評価し、emergent の算術能力を測る。
  • BIG-Bench の結果をメタ分析し、emergent 能力が指標依存であるかを評価する。
  • 計測指標を変えることでvisionタスクにおける誘発的 emergent 能力を示す。

提案手法

  • per-token クロスエントロピー損失 L_CE(N) がモデルサイズ N に対してべき乗則に従う単純な数学モデルを提供し、非線形または不連続な指標の下でタスクスコアへの非線形写像を導く。
  • 非線形指標(長い列における正解率など)や不連続な指標(Multiple Choice Grade など)は、小さなモデルから大きなモデルへの鋭い、emergent のような遷移を作り出す。
  • 連続的/線形指標(例:トークン編集距離、Brier スコア)はスケールとともに滑らかな改善をもたらし、emergent 効果を緩和する。
  • InstructGPT/GPT-3 での三つの予測を検証する:指標の変更で滑らかな改善が現れる;高解像度の統計で非ゼロの小モデル性能が非線形指標で示される;ターゲット長が予測可能に性能に影響する。
  • BIG-Bench の結果をメタ分析し、emergent 能力はごく限られた非線形/不連続な指標で主に現れ、連続的な指標では消失する。
  • 再構成タスクと逐次分類タスクの評価基準を設計することで、視覚モデルにも emergent 的能力を誘発できる。

実験結果

リサーチクエスチョン

  • RQ1 emergent abilities はモデル性能の評価に使用される指標に依存するか?
  • RQ2非線形または不連続な指標は、線形/連続な指標で消えるようなモデルスケールで鋭い遷移を生み出すか?
  • RQ3高解像度の統計(より多くのテストデータ)は、emergent 能力が主張されるタスクで小モデルの非ゼロの性能を明らかにするか?
  • RQ4評価指標を変更することで、非言語ドメイン(視覚)でも emergent 能力を誘発できるか?
  • RQ5BIG-Bench の emergent 主張は指標選択とモデルファミリーにどれくらい頑健か?

主な発見

  • emergent 能力は主に非線形または不連続な指標(Exact String Match や Multiple Choice Grade など)に現れる。
  • テストデータの解像度を高めると、小さなモデルに対しても上回る性能が現れ、非線形指標で非ゼロの能力が示される。
  • 線形または連続的な指標(例:Token Edit Distance や Brier Score)へ切り替えると、滑らかな、予測可能な改善が得られ、emergent 効果が薄れる。
  • メタ分析では、emergent 能力は指標のごく限られたサブセットでのみ現れ、2つの指標(Multiple Choice Grade と Exact String Match)が most of the claimed emergent abilities を説明する。
  • 指標を用いて、適切な評価基準を設計することで、アーキテクチャを跨ぐ視覚モデルに新たな emergent-like 能力を誘発できる。
  • InstructGPT/GPT-3 の経験的結果は、より良い統計と指標変更で emergent 能力が消えることを支持する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。