Skip to main content
QUICK REVIEW

[论文解读] Entropy and mutual information in models of deep neural networks

Marylou Gabrié, Andre Manoel|Infoscience (Ecole Polytechnique Fédérale de Lausanne)|May 24, 2018
Statistical Mechanics and Entropy被引用 14
一句话总结

该论文在正交不变权值矩阵的假设下,利用统计物理方法,提出了一种可计算的框架,用于在深度神经网络中计算熵和互信息。该框架通过自适应插值法,严格证明了在具有高斯权值的两层网络中,这些计算的正确性;并通过合成实验表明,在此设定下,压缩与泛化之间的关系仍不明确。

ABSTRACT

We examine a class of deep learning models with a tractable method to compute information-theoretic quantities. Our contributions are three-fold: (i) We show how entropies and mutual informations can be derived from heuristic statistical physics methods, under the assumption that weight matrices are independent and orthogonally-invariant. (ii) We extend particular cases in which this result is known to be rigorously exact by providing a proof for two-layers networks with Gaussian random weights, using the recently introduced adaptive interpolation method. (iii) We propose an experiment framework with generative models of synthetic datasets, on which we train deep neural networks with a weight constraint designed so that the assumption in (i) is verified during learning. We study the behavior of entropies and mutual informations throughout learning and conclude that, in the proposed setting, the relationship between compression and generalization remains elusive.

研究动机与目标

  • 开发一种可计算的方法,用于在深度神经网络中计算信息论量(如熵和互信息)。
  • 为深度学习模型中非线性激活和高维权值所使用的启发式统计物理公式建立严格的理论基础。
  • 通过受控的合成实验,研究压缩(通过层间互信息衡量)是否能解释深度网络中的泛化性能。
  • 通过在满足正交不变假设的约束权值矩阵上的生成模型进行数值实验,验证理论预测。

提出的方法

  • 使用具有逐分量非线性激活和 i.i.d. 噪声的随机多层前馈神经网络模型。
  • 应用统计物理中的副本方法,在正交不变权值矩阵的假设下,推导出熵和互信息的闭式表达式。
  • 采用自适应插值法,严格证明具有高斯权值的两层网络的副本公式。
  • 在训练过程中施加权值约束,以确保权值矩阵保持正交不变性,从而保证理论假设在整个学习过程中成立。
  • 构建具有受控相关性的合成数据集,并在此约束下训练深度网络,以研究信息流动。
  • 采用高维渐近分析,其中 $ n_0 \to \infty $,$ n_\ell / n_0 \to \alpha_\ell $,并利用大数集中不等式控制经验估计中的波动。

实验结果

研究问题

  • RQ1在非线性、高维设定下,是否能可靠地计算深度神经网络中的信息论量(如熵和互信息)?
  • RQ2副本方法对深度网络中互信息的预测,是否对具有高斯权值的两层模型严格成立?
  • RQ3在受控条件下,压缩(以层间互信息的减少衡量)是否与深度网络的泛化性能相关?
  • RQ4在理论假设成立的设定下,熵和互信息在训练过程中如何演化?

主要发现

  • 通过自适应插值法,严格证明了在具有 i.i.d. 高斯权值的两层网络中,深度网络的互信息和熵的副本公式成立。
  • 在具有约束权值矩阵的合成数据集上的数值实验表明,通过副本方法导出的理论预测具有高度准确性。
  • 尽管满足理论假设,训练过程中互信息的演化并未表现出清晰的压缩-泛化关系。
  • 研究发现,层间互信息在训练过程中并未一致减少,这挑战了压缩是泛化普遍驱动力的观点。
  • 通过集中不等式控制了经验互信息估计的方差,验证了数值框架的稳定性。
  • 该框架成功分离了权值矩阵结构的作用,表明正交不变权值使得信息论量的精确计算与验证成为可能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。