[论文解读] Hausdorff Dimension, Stochastic Differential Equations, and Generalization in Neural Networks.
本文通过将随机梯度下降(SGD)的轨迹建模为Feller过程,提出了一种深度学习中SGD的新型泛化界,表明泛化误差受这些轨迹的Hausdorff维数控制。关键结果表明,重尾噪声过程——以较低的尾指数为特征——可带来更好的泛化性能,从而确立尾指数作为衡量模型容量的新指标,该指标与测试误差相关,且不随模型规模增大而增长。
Despite its success in a wide range of applications, characterizing the generalization properties of stochastic gradient descent (SGD) in non-convex deep learning problems is still an important challenge. While modeling the trajectories of SGD via stochastic differential equations (SDE) under heavy-tailed gradient noise has recently shed light over several peculiar characteristics of SGD, a rigorous treatment of the generalization properties of such SDEs in a learning theoretical framework is still missing. Aiming to bridge this gap, in this paper, we prove generalization bounds for SGD under the assumption that its trajectories can be well-approximated by a Feller process, which defines a rich class of Markov processes that include several recent SDE representations (both Brownian or heavy-tailed) as its special case. We show that the generalization error can be controlled by the Hausdorff dimension of the trajectories, which is intimately linked to the tail behavior of the driving process. Our results imply that heavier-tailed processes should achieve better generalization; hence, the tail-index of the process can be used as a notion of ``capacity metric''. We support our theory with experiments on deep neural networks illustrating that the proposed capacity metric accurately estimates the generalization error, and it does not necessarily grow with the number of parameters unlike the existing capacity metrics in the literature.
研究动机与目标
- 为非凸深度学习中存在重尾梯度噪声的SGD,解决缺乏严格的理论泛化界的问题。
- 弥合SGD的随机微分方程(SDE)模型与泛化理论之间的差距。
- 基于SGD中驱动噪声过程的尾部行为,提出一种新的容量度量。
- 证明该度量与模型规模无关,仍能与泛化误差保持相关性。
提出的方法
- 将SGD轨迹建模为Feller过程,这是一类广义的马尔可夫过程,涵盖布朗运动和Lévy型SDE。
- 建立依赖于Feller过程轨迹Hausdorff维数的泛化界。
- 将Hausdorff维数与驱动噪声的尾部行为关联,表明更重的尾部(更低的尾指数)可降低泛化误差。
- 使用噪声过程的尾指数作为容量度量,替代传统的基于规模的度量方法。
- 在SGD动力学可被Feller过程良好近似的假设下,推导理论边界。
- 在深度神经网络上通过实验验证理论,将所提出的度量与实际泛化误差进行比较。
实验结果
研究问题
- RQ1能否在具有重尾噪声的随机微分方程框架下,严格推导出SGD的泛化界?
- RQ2SGD轨迹的Hausdorff维数与泛化性能之间有何关系?
- RQ3噪声过程的尾指数能否在深度学习中作为有意义的容量度量?
- RQ4所提出的容量度量是否与模型规模无关,仍能与泛化误差保持相关性?
- RQ5在实践中,该度量与现有容量度量相比表现如何?
主要发现
- SGD的泛化误差受其轨迹Hausdorff维数的限制,而该维数由驱动噪声过程的尾部行为决定。
- 尾部更重的噪声过程(尾指数更低)可导致更低的泛化误差,意味着更强的泛化能力。
- 噪声过程的尾指数可作为衡量泛化能力的新容量度量,且与泛化误差显著相关。
- 所提出的容量度量不会随参数数量的增加而必然增长,这与传统的参数数量或范数度量等不同。
- 在深度神经网络上的实证结果表明,该度量能准确预测不同架构和数据集下的泛化误差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。