Skip to main content
QUICK REVIEW

[論文レビュー] AI capabilities can be significantly improved without expensive retraining

Tom Davidson, Jean-Stanislas Denain|arXiv (Cornell University)|Dec 12, 2023
Adversarial Robustness in Machine Learning被引用数 6
ひとこと要約

本稿では、微調整やプロンプト設定、スキャフォールディングなどの後期学習AI向上手法の、同じ効果を得るために再訓練に要する計算資源と比較した際の性能向上を定量化するための「計算資源換算利得(CEG)」という指標を導入する。多くの向上手法が、再訓練に要する計算資源の5倍から20倍に相当する性能向上を達成しているが、そのコストはわずか1%未塔であることが判明した。これは、これらの手法がAI能力向上に極めて効率的な要因であることを示している。

ABSTRACT

State-of-the-art AI systems can be significantly improved without expensive retraining via "post-training enhancements"-techniques applied after initial training like fine-tuning the system to use a web browser. We review recent post-training enhancements, categorizing them into five types: tool-use, prompting methods, scaffolding, solution selection, and data generation. Different enhancements improve performance on different tasks, making it hard to compare their significance. So we translate improvements from different enhancements into a common currency, the compute-equivalent gain: how much additional training compute would be needed to improve performance by the same amount as the enhancement. Our non-experimental work shows that post-training enhancements have significant benefits: most surveyed enhancements improve benchmark performance by more than a 5x increase in training compute, some by more than 20x. Post-training enhancements are relatively cheap to develop: fine-tuning costs are typically <1% of the original training cost. Governing the development of capable post-training enhancements may be challenging because frontier models could be enhanced by a wide range of actors.

研究の動機と目的

  • 異なる手法間での性能向上を比較する課題を克服するため、共通の指標を用いて後期学習手法のAI能力への影響を定量化すること。
  • 後期学習手法が追加の前処理計算資源を費やすことと比較して、顕著な性能向上をもたらすかどうかを評価すること。
  • AI開発およびガバナンスにおける後期学習手法の経済的・戦略的影響を評価すること。
  • これらの手法がAI進歩のトレンドをどのように再編し、将来の安全性や政策意思決定にどのような影響を与えるかを検討すること。

提案手法

  • 計算資源換算利得(CEG)指標を導入:ある後期学習手法が達成する性能向上と同等の効果を得るために必要な追加の前処理計算量を定義する。
  • 後期学習手法を5つのタイプに分類する:ツール利用、プロンプト設定、スキャフォールディング、解決策選択、データ生成。
  • 同じ前処理分布上で計算量の増加による性能向上を仮定し、ベンチマークテスト前後のモデル性能を比較することでCEG値を推定する。
  • 元の前処理コストと比較して、後期学習手法の初期開発コストおよび実行時コストを評価し、コスト効率を測定する。
  • 既存の研究(例:MATH、HumanEval、GSM8K)のベンチマークデータを用いて、多様なタスクおよび手法タイプにおけるCEGを推定する。
  • CEGフレームワークを用いて、複数の後期学習手法の累積的および次第に減少する利得を評価し、スキルプロファイルおよび時間的改善速度を検討する。
Figure 1 : Illustration of the compute-equivalent gain. The enhanced model has the same performance as a non-enhanced model trained with $5\times$ more compute.
Figure 1 : Illustration of the compute-equivalent gain. The enhanced model has the same performance as a non-enhanced model trained with $5\times$ more compute.

実験結果

リサーチクエスチョン

  • RQ1後期学習手法がもたらす性能向上は、再訓練に要する計算資源と比較してどの程度のものか?
  • RQ2後期学習手法の開発および導入コストは、前処理計算量の増加と比較してどの程度か?
  • RQ3後期学習手法による能力向上は一般化可能なものか、それともタスク固有のものか?
  • RQ4複数の後期学習手法が相互にどのように作用するか。また、それらの組み合わせによる改善には、次第に減少する利得や上限があるのか?
  • RQ5後期学習による広範かつ低コストの能力向上が広がることで、AIガバナンスおよび安全性政策にどのような影響が生じるか?

主な発見

  • 調査された後期学習手法の多くは、前処理計算量を5倍以上増加させた場合と同等の性能向上を達成しており、一部の手法では20倍以上に相当する。
  • 多くの後期学習手法における微調整コストは、元の前処理コストの1%未塔であり、再訓練と比較して極めてコスト効率が良い。
  • 後期学習手法は、前処理スケーリングとは異なり、しばしばドメイン特異的である。
  • 計算資源換算利得(CEG)は、異なるベンチマークやドメインにおける多様な後期学習手法の影響を比較するうえで、有用で標準化された指標である。
  • 後期学習手法によるさらなる進歩の余地は大きく、しかし、次第に減少する利得や性能の上限が存在するかどうかは未解決のままである。
  • 後期学習手法の広範な利用可能性—多くの主体が容易に改善可能な可能性—は、AIガバナンスおよび安全性規制において、独自の課題を提起している。
Figure 2 : Distribution of CEG and additional costs of the techniques we studied. The one-time cost is given as a fraction of pre-training, the runtime cost is relative to the runtime cost without the enhancement. Enhancements without one-time cost are shown with an arrow on the y axis.
Figure 2 : Distribution of CEG and additional costs of the techniques we studied. The one-time cost is given as a fraction of pre-training, the runtime cost is relative to the runtime cost without the enhancement. Enhancements without one-time cost are shown with an arrow on the y axis.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。