Skip to main content
QUICK REVIEW

[论文解读] Towards a theory of machine learning

Vitaly Vanchurin|arXiv (Cornell University)|Apr 15, 2020
Statistical Mechanics and Entropy参考文献 45被引用 4
一句话总结

本文通过将神经网络建模为具有定义状态、权重、偏置和损失函数的七元组,提出了一种机器学习的统计力学框架。该框架通过最大熵和配分函数推导出学习的热力学定律,表明由于熵与复杂性之间的权衡,深度网络在最优学习效率下表现更佳,其结果对基础物理学中时空的涌现具有启示意义。

ABSTRACT

We define a neural network as a septuple consisting of (1) a state vector, (2) an input projection, (3) an output projection, (4) a weight matrix, (5) a bias vector, (6) an activation map and (7) a loss function. We argue that the loss function can be imposed either on the boundary (i.e. input and/or output neurons) or in the bulk (i.e. hidden neurons) for both supervised and unsupervised systems. We apply the principle of maximum entropy to derive a canonical ensemble of the state vectors subject to a constraint imposed on the bulk loss function by a Lagrange multiplier (or an inverse temperature parameter). We show that in an equilibrium the canonical partition function must be a product of two factors: a function of the temperature and a function of the bias vector and weight matrix. Consequently, the total Shannon entropy consists of two terms which represent respectively a thermodynamic entropy and a complexity of the neural network. We derive the first and second laws of learning: during learning the total entropy must decrease until the system reaches an equilibrium (i.e. the second law), and the increment in the loss function must be proportional to the increment in the thermodynamic entropy plus the increment in the complexity (i.e. the first law). We calculate the entropy destruction to show that the efficiency of learning is given by the Laplacian of the total free energy which is to be maximized in an optimal neural architecture, and explain why the optimization condition is better satisfied in a deep network with a large number of hidden layers. The key properties of the model are verified numerically by training a supervised feedforward neural network using the method of stochastic gradient descent. We also discuss a possibility that the entire universe on its most fundamental level is a neural network.

研究动机与目标

  • 通过统计力学原理,构建监督学习与无监督学习的统一理论框架。
  • 解决尽管深度学习具有高维参数空间,但其成功仍缺乏根本性解释的问题。
  • 定义一种适用于体部(隐藏层)而非仅边界(输入/输出神经元)的损失函数,从而实现无监督学习的形式化。
  • 推导神经网络学习过程的平衡态与非平衡态热力学定律。
  • 探讨宇宙本身可能本质上是受类似原理支配的神经网络的可能性。

提出的方法

  • 将神经网络定义为七元组:状态向量、输入/输出投影、权重矩阵、偏置向量、激活映射和损失函数。
  • 应用最大熵原理,通过拉格朗日乘子(逆温度)约束体部损失函数,推导出状态向量的系综。
  • 将系综配分函数计算为温度相关因子与权重和偏置函数的乘积,从而实现解析处理。
  • 推导学习的第一和第二定律:总熵减少至平衡态,且损失增量与热力学熵和复杂性增量成正比。
  • 引入熵破坏作为学习效率的度量,其与总自由能的拉普拉斯算子成正比。
  • 使用高斯积分近似与平滑窗函数处理有限激活范围,并定义算子 $\hat{G}$,其谱决定配分函数。

实验结果

研究问题

  • RQ1如何为监督与无监督系统构建统一的热力学描述?
  • RQ2在统计力学框架下,体部(隐藏层)损失函数在实现无监督学习中起什么作用?
  • RQ3由最大熵导出的系综如何与神经网络的平衡态相关联?
  • RQ4根据推导出的热力学模型,为何深层网络(具有大量隐藏层)的学习效率更高?
  • RQ5能否从神经网络非平衡态热力学中推导出时空与广义相对论的涌现?

主要发现

  • 系综配分函数可分解为温度相关项与依赖权重和偏置的项,表明总自由能可分解为热力学与复杂性两部分。
  • 总香农熵可分解为热力学熵与复杂性项,二者均参与学习过程。
  • 学习的第一定律指出:损失的变化与热力学熵和网络复杂性变化之和成正比。
  • 学习的第二定律成立:训练过程中总熵持续减少,直至达到平衡态。
  • 当总自由能的拉普拉斯算子最大时,学习效率达到峰值,这有利于具有大量隐藏层的深层架构。
  • 在对昂萨格张量作特定对称性假设下,熵产生可导出爱因斯坦场方程,表明广义相对论可能在大尺度上从神经网络动力学中涌现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。