Skip to main content
QUICK REVIEW

[論文レビュー] Large-time asymptotics in deep learning

Carlos Esteve, Borjan Geshkovski|arXiv (Cornell University)|Aug 6, 2020
Model Reduction and Neural Networks参考文献 93被引用数 9
ひとこと要約

この論文は、ニューラルODEフレームワークを用いて深層学習の長時間漈態を研究し、トレーニング誤差が最終時間 $T$ に対して $\mathcal{O}(1/T)$ で減少すること、最適パラメータが最小 $L^2$-ノルム補間関数に収束することを示している。ターンパイク理論にインspiredされた積分正則化項を導入することで、誤差およびパラメータの両方が指数的減少 $\mathcal{O}(e^{-\mu t})$ を示し、可変幅のResNetsを用いた連続的空間時間ニューラルネットワークへの拡張も可能となる。

ABSTRACT

We consider the neural ODE perspective of supervised learning and study the impact of the final time $T$ (which may indicate the depth of a corresponding ResNet) in training. For the classical $L^2$--regularized empirical risk minimization problem, whenever the neural ODE dynamics are homogeneous with respect to the parameters, we show that the training error is at most of the order $\mathcal{O}\left(\frac{1}{T} ight)$. Furthermore, if the loss inducing the empirical risk attains its minimum, the optimal parameters converge to minimal $L^2$--norm parameters which interpolate the dataset. By a natural scaling between $T$ and the regularization hyperparameter $λ$ we obtain the same results when $λ\searrow0$ and $T$ is fixed. This allows us to stipulate generalization properties in the overparametrized regime, now seen from the large depth, neural ODE perspective. To enhance the polynomial decay, inspired by turnpike theory in optimal control, we propose a learning problem with an additional integral regularization term of the neural ODE trajectory over $[0,T]$. In the setting of $\ell^p$--distance losses, we prove that both the training error and the optimal parameters are at most of the order $\mathcal{O}\left(e^{-μt} ight)$ in any $t\in[0,T]$. The aforementioned stability estimates are also shown for continuous space-time neural networks, taking the form of nonlinear integro-differential equations. By using a time-dependent moving grid for discretizing the spatial variable, we demonstrate that these equations provide a framework for addressing ResNets with variable widths.

研究の動機と目的

  • 最終時間 $T$ がニューラルODEの視点におけるトレーニング誤差および一般化に与える影響を分析すること。
  • 過パラメータ化された状態において $T \to \infty$ のとき、最適パラメータが最小 $L^2$-ノルム補間関数に収束することを確立すること。
  • 最適制御理論のターンパイク理論にインspiredされた積分正則化に基づく拡張された経験的リスクを用いて、多項式的減少率の改善を図ること。
  • 非線形積分微分方程式に従う連続的空間時間ニューラルネットワークへの結果の拡張。
  • 可変幅のResNetsが連続的空間時間定式化における時間依存移動グリッドを用いてモデル化可能であることを示すこと。

提案手法

  • 深さを時間ホライズンとみなして、最終時間 $T$ を用いたニューラルODE最適制御問題として教師あり学習を定式化する。
  • 均一なダイナミクス下での $L^2$-正則化された経験的リスク最小化を分析し、$\mathcal{O}(1/T)$ のトレーニング誤差減少を導出する。
  • 軌道上の $[0,T]$ における積分正則化項を追加した拡張損失を導入し、ターンパイク的挙動と指数的安定性を強制する。
  • グローワルの不等式と可制御性の議論を用いて、軌道およびパラメータの安定性推定を導出する。
  • 時間依存移動グリッドを用いて連続的空間時間ニューラルネットワークの空間変数を離散化し、可変幅のResNetsのモデル化を可能にする。
  • 拡張定式化下で $\ell^p$-損失に対して、トレーニング誤差および最適パラメータの両方が $\mathcal{O}(e^{-\mu t})$ で指数的減少することを証明する。

実験結果

リサーチクエスチョン

  • RQ1ニューラルODEにおける最終時間 $T$ を大きくすると、$L^2$-正則化された学習におけるトレーニング誤差にどのような影響を与えるか?
  • RQ2過パラメータ化された状態において $T \to \infty$ のとき、最適パラメータの極限的挙動は何か?
  • RQ3軌道全体にわたる積分正則化を用いることで、$\mathcal{O}(1/T)$ よりも速い収束速度を達成できるか?
  • RQ4可変幅のニューラルネットワークに対して、連続的空間時間ニューラルネットワークへの結果の拡張はどのように行われるか?
  • RQ5ターンパイク理論は、ニューラルODE学習における指数的収束を達成するために果たす役割は何か?

主な発見

  • パラメータに依存しない均一なダイナミクス下では、$L^2$-正則化された経験的リスク最小化において、トレーニング誤差は $\mathcal{O}(1/T)$ で減少する。
  • 最適パラメータは $T \to \infty$ のとき、データセットの最小 $L^2$-ノルム補間関数に収束し、過パラメータ化された状態で一般化を保証する。
  • パラメータ $\lambda \searrow 0$ をスケーリングし、$T$ を固定した場合でも、同じ $\mathcal{O}(1/T)$ の減少率が維持され、深さと正則化の関係が明確化される。
  • 追加の積分正則化項を導入することで、$t \in [0,T]$ 全体にわたり、トレーニング誤差および最適パラメータの両方が指数的減少 $\mathcal{O}(e^{-\mu t})$ を示す。
  • 指数的減少の結果は $\ell^p$-距離損失に対しても成り立ち、非線形積分微分方程式に従う連続的空間時間ニューラルネットワークへも拡張可能である。
  • 時間依存移動グリッドを用いることで、連続的空間時間フレームワークが可変幅のResNetsをモデル化可能となり、安定性および収束性の性質が保持される。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。