Skip to main content
QUICK REVIEW

[论文解读] Generalization Error Bounds Via Rényi-, $f$-Divergences and Maximal Leakage

Amedeo Roberto Esposito, Michael Gastpar|arXiv (Cornell University)|Dec 1, 2019
Distributed Sensor Networks and Detection Algorithms参考文献 30被引用 15
一句话总结

本文通过利用 Rényi 散度、f-散度和最大泄漏(maximal leakage),推导了依赖随机变量的一般化误差界,将经典的不等式(如 Hoeffding 和 McDiarmid)扩展至自适应和依赖设定。主要贡献是基于最大泄漏的鲁棒界,该界在自适应数据分析中具有良好的组合性,并可推广至非 i.i.d. 场景。

ABSTRACT

In this work, the probability of an event under some joint distribution is bounded by measuring it with the product of the marginals instead (which is typically easier to analyze) together with a measure of the dependence between the two random variables. These results find applications in adaptive data analysis, where multiple dependencies are introduced and in learning theory, where they can be employed to bound the generalization error of a learning algorithm. Bounds are given in terms of Sibson's Mutual Information, $α-$Divergences, Hellinger Divergences, and $f-$Divergences. A case of particular interest is the Maximal Leakage (or Sibson's Mutual Information of order infinity), since this measure is robust to post-processing and composes adaptively. The corresponding bound can be seen as a generalization of classical bounds, such as Hoeffding's and McDiarmid's inequalities, to the case of dependent random variables.

研究动机与目标

  • 为自适应数据分析中的泛化误差提供边界,其中查询之间的依赖关系源于顺序查询。
  • 通过使用散度度量依赖性,将经典集中不等式(如 Hoeffding 和 McDiarmid)扩展至具有依赖随机变量的设定。
  • 开发对后处理鲁棒且可自适应组合的边界,这对于分析自适应数据分析中的机制至关重要。
  • 利用 Luxemburg 和 Amemiya 范数,将各种基于散度的边界(Rényi、f-散度、Hellinger)统一于单一理论框架下。
  • 建立信息论依赖度量与学习算法中泛化误差之间的联系。

提出的方法

  • 推导出形式为 $\mathcal{P}(E) \leq f(\mathcal{Q}(E)) \cdot g(d\mathcal{P}/d\mathcal{Q})$ 的一般边界,其中 $\mathcal{P}$ 为联合分布,$\mathcal{Q}$ 为各边缘分布的乘积。
  • 应用 Luxemburg 和 Amemiya 范数,基于 f-散度和 Rényi 散度推导边界,实现对依赖性的灵活量化。
  • 使用 Sibson 的 α 阶互信息作为依赖度量,特别关注 $\alpha \to \infty$ 的极限情形,对应于最大泄漏。
  • 采用共轭凸函数方法与对偶性,推导涉及 Hellinger 散度和 f-互信息的边界。
  • 证明最大泄漏边界在独立性条件下退化为经典不等式,验证其普遍性。
  • 建立最大泄漏对后处理鲁棒且可自适应组合的性质,使其特别适用于自适应数据分析。

实验结果

研究问题

  • RQ1当涉及的随机变量为依赖关系而非 i.i.d. 时,如何对泛化误差进行边界界定?
  • RQ2Rényi 散度和 $f$-散度能否用于推导更紧致的泛化边界,以反映学习算法中的统计依赖性?
  • RQ3依赖度量需满足何种性质才能在自适应数据分析中有用?最大泄漏是否满足这些标准?
  • RQ4所提出的框架如何将 Hoeffding 和 McDiarmid 等经典集中不等式推广至依赖设定?
  • RQ5基于最大泄漏的边界是否可高效计算,并可应用于添加噪声的实际学习机制?

主要发现

  • 本文推导出基于 α > 1 阶 Rényi 散度的一般边界,表明 $\mathcal{P}_{XY}(E) \leq \mathcal{P}_X\mathcal{P}_Y(E)^{\frac{\alpha-1}{\alpha}} \cdot \exp\left(\frac{\alpha-1}{\alpha} D_\alpha(\mathcal{P}_{XY} \| \mathcal{P}_X\mathcal{P}_Y)\right)$。
  • 对于最大泄漏($\alpha \to \infty$),边界简化为 $\mathcal{P}_{XY}(E) \leq \mathcal{P}_X\mathcal{P}_Y(E) \cdot \exp\left(\mathcal{L}(X \to Y)\right)$,该边界具有鲁棒性且可自适应组合。
  • 通过 $\phi$-函数 $\phi_\alpha(t) = \frac{t^\alpha - 1}{\alpha - 1}$ 的共轭函数,推导出基于 Hellinger 散度 α 阶的边界,得到与散度呈类似指数依赖的形式。
  • 证明最大泄漏可通过 $\mathcal{L}(X \to Y) = \log \sum_y \max_{x: P(x)>0} P_{Y|X}(y|x)$ 计算,从而可实际应用于带噪声的学习算法。
  • 当 $\mathcal{P}_{XY} = \mathcal{P}_X\mathcal{P}_Y$ 时,边界退化为经典集中不等式,验证了其一致性。
  • 该框架具有一般性,适用于任意满足绝对连续性的概率测度对,不限于联合与乘积测度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。