Skip to main content
QUICK REVIEW

[论文解读] Tighter expected generalization error bounds via Wasserstein distance

Borja Rodríguez-Gálvez, Germán Bassi|arXiv (Cornell University)|Jan 22, 2021
Adversarial Robustness in Machine Learning参考文献 13被引用 20
一句话总结

本文通过Wasserstein距离引入了更紧致的期望泛化误差界,提出了涵盖全数据集、单字母形式及随机子集的边界,其性能优于现有的基于相对熵的边界。通过利用Wasserstein度量对假设空间几何结构的建模,这些边界非平凡,弥合了几何方法与信息论方法在泛化分析中的鸿沟。

ABSTRACT

This work presents several expected generalization error bounds based on the Wasserstein distance. More specifically, it introduces full-dataset, single-letter, and random-subset bounds, and their analogues in the randomized subsample setting from Steinke and Zakynthinou [1]. Moreover, when the loss function is bounded and the geometry of the space is ignored by the choice of the metric in the Wasserstein distance, these bounds recover from below (and thus, are tighter than) current bounds based on the relative entropy. In particular, they generate new, non-vacuous bounds based on the relative entropy. Therefore, these results can be seen as a bridge between works that account for the geometry of the hypothesis space and those based on the relative entropy, which is agnostic to such geometry. Furthermore, it is shown how to produce various new bounds based on different information measures (e.g., the lautum information or several $f$-divergences) based on these bounds and how to derive similar bounds with respect to the backward channel using the presented proof techniques.

研究动机与目标

  • 开发能反映假设空间几何结构的更紧致的期望泛化误差界。
  • 解决基于相对熵的边界在互信息发散时可能变得平凡(vacuous)的局限性。
  • 通过Wasserstein距离统一信息论泛化界与几何洞察。
  • 将现有的随机子样本框架扩展,以整合基于Wasserstein的依赖关系,实现有限且非平凡的边界。

提出的方法

  • 基于条件分布之间的Wasserstein距离,提出全数据集、单字母形式及随机子集的泛化误差界。
  • 利用Kantorovich-Rubinstein对偶性,将期望泛化误差表示为假设分布与其边缘分布之间Wasserstein距离的形式。
  • 通过联合范围策略及f-散度(包括卡方散度)的变分表示推导边界。
  • 将该框架应用于标准设置与随机子样本设置,通过条件互信息项实现控制。
  • 提出一种新颖的证明技术,使基于Lautum信息与f-散度等替代信息度量的边界推导成为可能。
  • 展示如何利用后向信道依赖关系,在相同理论框架下推导出类似边界。

实验结果

研究问题

  • RQ1基于Wasserstein距离的边界是否能实现比现有基于相对熵的边界更紧致的泛化误差估计?
  • RQ2当显式建模假设空间几何结构时,Wasserstein边界的表现如何?
  • RQ3能否通过在相对熵框架中引入基于Wasserstein距离的几何结构,导出非平凡的泛化边界?
  • RQ4在随机子样本设置中,Wasserstein边界与条件互信息之间存在何种关系?
  • RQ5所提出的框架如何扩展至其他信息度量(如f-散度与Lautum信息)?

主要发现

  • 所提出的基于Wasserstein的边界比现有基于相对熵的边界更紧致,并在忽略几何结构时可恢复为后者。
  • 当损失函数有界时,即使在相对熵边界变得平凡的情况下,该边界仍能恢复出非平凡的估计。
  • 卡方散度的变分表示所导出的边界,比通过联合范围策略导出的边界更紧致且更具普适性。
  • 当损失方差较小时,且卡方散度低于约2.51时,基于卡方散度的边界(公式15)比基于Pinsker的边界(公式5)更紧致。
  • 该框架通过统一的证明技术,可推导出基于Lautum信息与各类f-散度等替代信息度量的新边界。
  • 条件互信息项如$ I(W;U| ilde{S}) $在连续设置中保持有限且小于无条件对应项,确保了鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。