Skip to main content
QUICK REVIEW

[论文解读] Generalization Bounds via Convex Analysis

Gábor Lugosi, Gergely Neu|arXiv (Cornell University)|Feb 10, 2022
Sparse and Compressive Sensing Techniques被引用 6
一句话总结

本文通过用任意强凸依赖度量替代香农互信息,推广了信息论一般化边界,利用凸分析追踪势函数。它建立了基于p-范数散度和Wasserstein-2距离的新边界,适用于重尾损失和光滑损失函数,在由适当范数捕获的广义子高斯型尾部条件下适用。

ABSTRACT

Since the celebrated works of Russo and Zou (2016,2019) and Xu and Raginsky (2017), it has been well known that the generalization error of supervised learning algorithms can be bounded in terms of the mutual information between their input and the output, given that the loss of any fixed hypothesis has a subgaussian tail. In this work, we generalize this result beyond the standard choice of Shannon's mutual information to measure the dependence between the input and the output. Our main result shows that it is indeed possible to replace the mutual information by any strongly convex function of the joint input-output distribution, with the subgaussianity condition on the losses replaced by a bound on an appropriately chosen norm capturing the geometry of the dependence measure. This allows us to derive a range of generalization bounds that are either entirely new or strengthen previously known ones. Examples include bounds stated in terms of $p$-norm divergences and the Wasserstein-2 distance, which are respectively applicable for heavy-tailed loss distributions and highly smooth loss functions. Our analysis is entirely based on elementary tools from convex analysis by tracking the growth of a potential function associated with the dependence measure and the loss function.

研究动机与目标

  • 将现有的信息论一般化边界从香农互信息推广至更广范围。
  • 利用凸分析建立统一框架,用于分析监督学习中的依赖度量。
  • 用基于范数的条件替代子高斯损失假设,该条件与所选依赖度量相适应。
  • 为多样化损失分布(包括重尾和光滑情形)推导新的或改进的一般化边界。
  • 提供一种系统化方法,通过凸性约束下的势函数增长推导边界。

提出的方法

  • 该框架使用输入-输出联合分布与乘积分布之间强凸散度度量的势函数。
  • 利用凸分析工具追踪该势函数在数据点上的增长。
  • 该方法用基于范数的条件替代子高斯尾部假设,以反映所选依赖度量的几何结构。
  • 适用于任意强凸函数形式的联合分布,超越互信息的推广。
  • 分析基于基础凸分析,避免使用复杂的概率工具。
  • 通过将期望损失差与散度增长及范数约束关联,推导出一般化边界。

实验结果

研究问题

  • RQ1能否使用除香农互信息外的其他依赖度量推导一般化边界?
  • RQ2凸分析如何用于分析任意强凸散度相关势函数的增长?
  • RQ3当使用p-范数散度或Wasserstein-2距离作为依赖度量时,损失函数需要何种尾部条件?
  • RQ4该框架能否为重尾损失分布或高度光滑损失函数提供更紧或全新的边界?
  • RQ5在不同依赖度量下,范数几何结构在塑造泛化误差边界中起什么作用?

主要发现

  • 本文建立了使用任意强凸依赖度量的一般化边界,不仅限于互信息。
  • 它用依赖于所选散度的基于范数的条件替代了子高斯损失假设,增强了适用范围。
  • 推导出基于p-范数散度的新边界,对重尾损失分布具有有效性。
  • 证明Wasserstein-2距离可为高度光滑损失函数提供紧致边界。
  • 该框架通过系统性地将依赖度量的选择与相应损失尾部条件通过凸势函数关联,统一并强化了先前结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。