[论文解读] An Optimal Transport View on Generalization
本文提出了一种新颖的最优传输框架,用于在机器学习中界定泛化误差,利用算法传输成本——即模型输出与其给定训练样本的条件输出之间的期望Wasserstein距离。该框架在不假设子高斯或有界损失的前提下推导出边界,并将其与信息论和学习理论概念(如VC维和KL散度)联系起来,表明深层神经网络由于分层结构和f-散度压缩,其泛化误差随深度增加呈指数衰减。
We derive upper bounds on the generalization error of learning algorithms based on their \emph{algorithmic transport cost}: the expected Wasserstein distance between the output hypothesis and the output hypothesis conditioned on an input example. The bounds provide a novel approach to study the generalization of learning algorithms from an optimal transport view and impose less constraints on the loss function, such as sub-gaussian or bounded. We further provide several upper bounds on the algorithmic transport cost in terms of total variation distance, relative entropy (or KL-divergence), and VC dimension, thus further bridging optimal transport theory and information theory with statistical learning theory. Moreover, we also study different conditions for loss functions under which the generalization error of a learning algorithm can be upper bounded by different probability metrics between distributions relating to the output hypothesis and/or the input data. Finally, under our established framework, we analyze the generalization in deep learning and conclude that the generalization error in deep neural networks (DNNs) decreases exponentially to zero as the number of layers increases. Our analyses of generalization error in deep learning mainly exploit the hierarchical structure in DNNs and the contraction property of $f$-divergence, which may be of independent interest in analyzing other learning models with hierarchical structure.
研究动机与目标
- 开发一种基于最优传输理论的新理论框架,用于分析学习算法中的泛化误差。
- 推导出不依赖于损失函数严格假设(如子高斯性或有界性)的泛化误差边界。
- 通过将算法传输成本与总变差、KL散度和VC维等度量关联,建立最优传输与信息论及统计学习理论之间的联系。
- 分析深层神经网络(DNNs)中的泛化行为,并解释为何其在高容量下仍能良好泛化。
- 确立DNN中泛化误差随深度呈指数衰减的结论,归因于分层结构和f-散度的压缩特性。
提出的方法
- 将算法传输成本定义为模型输出假设与其给定训练样本的条件版本之间的期望Wasserstein距离。
- 基于此传输成本推导出泛化误差的上界,适用于利普希茨连续损失函数且无需分布假设。
- 通过不等式将算法传输成本与其它概率度量关联,包括总变差、相对熵和赫林格距离。
- 通过VC维及其他复杂性度量对传输成本进行边界估计,建立与经典学习理论的联系。
- 将该框架应用于深度学习,将DNN建模为分层特征映射的马尔可夫链,并利用子高斯散度不等式(SDPI)对各层之间的互信息进行边界估计。
- 利用f-散度在各层间的压缩特性,证明最终假设与训练数据之间的互信息随深度呈指数衰减。
实验结果
研究问题
- RQ1能否在不假设子高斯或有界损失函数的前提下,利用最优传输度量界定泛化误差?
- RQ2算法传输成本如何与经典学习理论概念(如VC维)及信息论度量(如KL散度)关联?
- RQ3深层神经网络的分层结构在控制泛化误差方面发挥何种作用?
- RQ4f-散度的压缩特性是否可用于推导深度学习中泛化误差的指数衰减?
- RQ5所提出的框架如何通过正则化实现高概率泛化边界和算法设计?
主要发现
- 泛化误差的上界由输出假设与其给定训练样本的条件版本之间的期望Wasserstein距离决定,适用于利普希茨损失函数且无需分布假设。
- 算法传输成本可通过总变差距离、相对熵和VC维进行界定,从而在最优传输与信息论及学习理论之间建立桥梁。
- 对于有界损失函数,推导出一种类似总变差的泛化边界,该边界可通过度量不等式进一步由赫林格距离和χ²距离界定。
- 在深层神经网络中,由于分层结构和f-散度在各层间的压缩,泛化误差随层数增加呈指数衰减。
- 最终假设与训练数据之间的互信息随深度呈指数衰减,衰减速率由各层间压缩系数的几何平均值决定。
- 该框架表明,可通过将推导出的泛化边界纳入目标函数来设计正则化方法,以平衡拟合与泛化性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。