Skip to main content
QUICK REVIEW

[論文レビュー] Towards a theory of machine learning

Vitaly Vanchurin|arXiv (Cornell University)|Apr 15, 2020
Statistical Mechanics and Entropy参考文献 45被引用数 4
ひとこと要約

本稿は、ニューラルネットワークを状態、重み、バイアス、損失関数を明確に定義したセプティプレットとしてモデル化することにより、機械学習の統計力学的枠組みを提案する。最大エントロピーと分配関数を用いて学習の熱力学的法則を導出し、最適な学習効率がエントロピーと複雑性のトレードオフによって深層ネットワークで生じることを示し、素因数物理学における出現的時空にインパクトを与える。

ABSTRACT

We define a neural network as a septuple consisting of (1) a state vector, (2) an input projection, (3) an output projection, (4) a weight matrix, (5) a bias vector, (6) an activation map and (7) a loss function. We argue that the loss function can be imposed either on the boundary (i.e. input and/or output neurons) or in the bulk (i.e. hidden neurons) for both supervised and unsupervised systems. We apply the principle of maximum entropy to derive a canonical ensemble of the state vectors subject to a constraint imposed on the bulk loss function by a Lagrange multiplier (or an inverse temperature parameter). We show that in an equilibrium the canonical partition function must be a product of two factors: a function of the temperature and a function of the bias vector and weight matrix. Consequently, the total Shannon entropy consists of two terms which represent respectively a thermodynamic entropy and a complexity of the neural network. We derive the first and second laws of learning: during learning the total entropy must decrease until the system reaches an equilibrium (i.e. the second law), and the increment in the loss function must be proportional to the increment in the thermodynamic entropy plus the increment in the complexity (i.e. the first law). We calculate the entropy destruction to show that the efficiency of learning is given by the Laplacian of the total free energy which is to be maximized in an optimal neural architecture, and explain why the optimization condition is better satisfied in a deep network with a large number of hidden layers. The key properties of the model are verified numerically by training a supervised feedforward neural network using the method of stochastic gradient descent. We also discuss a possibility that the entire universe on its most fundamental level is a neural network.

研究の動機と目的

  • 統計力学の原則を用いて教師あり学習と教師なし学習の両方を統一的に記述する理論的枠組みを構築すること。
  • 高次元パラメータ空間を持つにもかかわらず深層学習の成功に根本的な説明が欠けているという問題に取り組むこと。
  • 境界(入力/出力ニューロン)でのみ適用可能な損失関数ではなく、バルク(隠れ層)における適用が可能な損失関数を定義し、教師なし学習の形式的記述を可能にすること。
  • ニューラルネットワーク学習プロセスの平衡および非平衡熱力学的法則を導出すること。
  • 宇宙そのものが同様の原理に従うニューラルネットワークである可能性を検討すること。

提案手法

  • ニューラルネットワークを、状態ベクトル、入出力射影、重み行列、バイアスベクトル、活性化マップ、損失関数からなるセプティプレットとして定義する。
  • ラグランジュ乗数(逆温度)を用いてバルク損失関数によって制約された状態ベクトルのカノニカル集合を最大エントロピーの原理によって導出する。
  • 温度依存係数と重み・バイアスに依存する関数の積として分配関数を計算し、解析的取り扱いを可能にする。
  • 学習の第一法則および第二法則を導出:全エントロピーは平衡に達するまで減少し、損失の増分は熱力学的エントロピーと複雑性増分の和に比例する。
  • 全自由エネルギーのラプラシアンに比例するエントロピー破壊を学習効率の指標として導入する。
  • 有限な活性化範囲を扱うためにガウス積分近似と滑らかな窓関数を用い、分配関数を決定する作用素 $\hat{G}$ を定義する。

実験結果

リサーチクエスチョン

  • RQ1教師ありおよび教師なしシステムの両方に対して統一的な熱力学的記述を構築するにはどうすればよいか?
  • RQ2統計力学的枠組み内での教師なし学習を可能にするために、バルク(隠れ層)損失関数が果たす役割は何か?
  • RQ3最大エントロピーから導出されたカノニカル集合は、ニューラルネットワークの平衡状態とどのように関係するか?
  • RQ4統計力学モデルの観点から、なぜ多くの隠れ層を持つ深層ネットワークで学習効率が高まるのか?
  • RQ5非平衡熱力学的性質に基づいて、時空や一般相対性理論の出現をニューラルネットワークのダイナミクスから導出できるか?

主な発見

  • カノニカル分配関数は温度依存項と重み・バイアスに依存する項に分解され、全自由エネルギーが熱力学的および複雑性的成分に分解されることを示唆する。
  • 全シャノンエントロピーは熱力学的エントロピーと複雑性項に分解され、両者とも学習プロセスに寄与する。
  • 学習の第一法則は、損失の変化が熱力学的エントロピーの変化とネットワーク複雑性の変化の和に比例することを示す。
  • 学習の第二法則は、全エントロピーが平衡に達するまで減少することを保証する。
  • 学習効率は、全自由エネルギーのラプラシアンが最大になるときに最大化され、多くの隠れ層を持つ深層アーキテクチャが好ましい。
  • オンサッガー張子に特定の対称性を仮定すると、エントロピー生成はアインシュタインの場の運動方程式へと導かれる。これは、一般相対性理論が大規模スケールでニューラルネットワークのダイナミクスから出現しうることを示唆する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。