Skip to main content
QUICK REVIEW

[論文レビュー] Identifying Equivalent Training Dynamics

William T. Redman, Juan M. Bello‐Rivas|arXiv (Cornell University)|Feb 17, 2023
Model Reduction and Neural Networks被引用数 6
ひとこと要約

本稿では、最適化軌道のスペクトル類似性を分析することにより、機械学習における等価な学習ダイナミクスを同定するデータ駆動型フレームワークを、Koopman作用素理論を用いて提案する。実証的に、学習率とバッチサイズの比、層の幅、データセットの種別、活性化関数が、順伝播ニューラルネットワークにおける共役性を規定することを明らかにした。

ABSTRACT

Study of the nonlinear evolution deep neural network (DNN) parameters undergo during training has uncovered regimes of distinct dynamical behavior. While a detailed understanding of these phenomena has the potential to advance improvements in training efficiency and robustness, the lack of methods for identifying when DNN models have equivalent dynamics limits the insight that can be gained from prior work. Topological conjugacy, a notion from dynamical systems theory, provides a precise definition of dynamical equivalence, offering a possible route to address this need. However, topological conjugacies have historically been challenging to compute. By leveraging advances in Koopman operator theory, we develop a framework for identifying conjugate and non-conjugate training dynamics. To validate our approach, we demonstrate that comparing Koopman eigenvalues can correctly identify a known equivalence between online mirror descent and online gradient descent. We then utilize our approach to: (a) identify non-conjugate training dynamics between shallow and wide fully connected neural networks; (b) characterize the early phase of training dynamics in convolutional neural networks; (c) uncover non-conjugate training dynamics in Transformers that do and do not undergo grokking. Our results, across a range of DNN architectures, illustrate the flexibility of our framework and highlight its potential for shedding new light on training dynamics.

研究の動機と目的

  • 機械学習の訓練における異なるハイパーパrameter設定が等価な最適化軌道をもたらす条件を特定するための一般化された手法の欠如に対処すること。
  • 特に共役性の観点から、最適化アルゴリズムの背後にある力学的挙動に基づいて分類を可能にすること。
  • 内部のアルゴリズム的式にアクセスできない状況でも、多様なML手法間での訓練ダイナミクスを比較可能なデータ駆動型で一般化可能なアプローチを提供すること。
  • Koopman作用素理論を、特に深層ニューラルネットワークにおける反復的ML最適化プロセスの分析に拡張すること。
  • 学習率、バッチサイズ、幅、活性化関数、データセットといったハイパーパrameterが、訓練中の共役または非共役ダイナミクスを引き起こす要因となるかを解明すること。

提案手法

  • 各更新ステップが状態遷移写像に対応する離散時間力学系として、反復的ML最適化をモデル化する。
  • 観測された訓練軌道からスペクトル的対象(固有値とモード)を抽出するために、Koopmanモード分解を適用する。
  • スペクトル類似性、特に重複するKoopman固有値を基準として、共役力学系(等価な最適化ダイナミクスを示す)を同定する。
  • オンラインミラー降下法と勾配降下法の比較に、データ駆動型Koopman解析を用い、解析的に示された等価性をデータドリブンな形で確認する。
  • 全結合ネットワークにおいて、学習率、バッチサイズ、幅、活性化関数、データセットといったハイパーパrameterを体系的に変化させ、Koopmanスペクトルに与える影響を評価する。
  • 複数のランダムシードにわたるKoopmanスペクトル間のWasserstein距離を用いて、スペクトル類似性の統計的安定性を定量的に評価する。
Figure 1: Data-driven identification of the conjugacy between online mirror and gradient descent. (A) Example trajectory for $x_{t}$ (red) and $q(u_{t})$ (black), when $f(x)=\sum_{i=1}^{d}x_{i}^{4}$ . Rhombus denotes initialization and star denotes the end of $500$ iterations. (B) Koopman spectra co
Figure 1: Data-driven identification of the conjugacy between online mirror and gradient descent. (A) Example trajectory for $x_{t}$ (red) and $q(u_{t})$ (black), when $f(x)=\sum_{i=1}^{d}x_{i}^{4}$ . Rhombus denotes initialization and star denotes the end of $500$ iterations. (B) Koopman spectra co

実験結果

リサーチクエスチョン

  • RQ1深層ニューラルネットワークの訓練において、異なるハイパーパラメータ設定が等価な最適化ダイナミクスをもたらす条件は何か?
  • RQ2Koopmanスペクトル類似性は、オンラインミラー降下法と勾配降下法といった異なる最適化アルゴリズム間の共役性を信頼性高く検出できるか?
  • RQ3学習率、バッチサイズ、層幅、データセット、活性化関数といったハイパーパラメータのうち、訓練ダイナミクスの共役構造に最も顕著に影響を与えるのはどれか?
  • RQ4データ駆動型Koopman解析は、内部のアルゴリズム的式にアクセスできない状況でも、等価な訓練軌道をどれほど正確に同定できるか?
  • RQ5MNISTや合成データなど、異なるデータセットや活性化関数は、最適化プロセスのスペクトル構造と共役性にどのように影響を与えるか?

主な発見

  • 特定の問題において、オンラインミラー降下法と勾配降下法の間で、Koopmanスペクトルが顕著に重複しており、解析的に示された等価性を裏付ける。
  • 学習率とバッチサイズの比が、最適化ダイナミクスが共役であるかどうかを決定づける重要な要因である。
  • 層の幅とデータセットの性質(例:手書き文字 vs. 合成データ)が、訓練ダイナミクスの共役構造に顕著に影響を与える。
  • 活性化関数の選択が、Koopman作用素のスペクトル的性質に影響を与え、結果として最適化軌道の等価性に影響を及ぼす。
  • Koopmanモード分解によるスペクトル解析から、同じ訓練性能を示す場合でも、異なるハイパーパラメータの組み合わせが非共役ダイナミクスをもたらすことが明らかになった。
  • 本フレームワークは、根の探索アルゴリズムやニューラルネットワーク訓練を含む多様な状況において共役性を的確に同定でき、広範な適用可能性を示した。
Figure 2: Learning rate to batch size ratio controls the conjugacy of training. (A) Training loss, as a function of number of images shown, for different choices of batch size. Loss was first smoothed using a square filter. Width of curves denotes $25^{\text{th}}$ and $75^{\text{th}}$ percentile of
Figure 2: Learning rate to batch size ratio controls the conjugacy of training. (A) Training loss, as a function of number of images shown, for different choices of batch size. Loss was first smoothed using a square filter. Width of curves denotes $25^{\text{th}}$ and $75^{\text{th}}$ percentile of

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。