Skip to main content
QUICK REVIEW

[论文解读] $(f,Γ)$-Divergences: Interpolating between $f$-Divergences and Integral Probability Metrics

Jeremiah Birrell, Paul Dupuis|arXiv (Cornell University)|Nov 11, 2020
Mathematical Analysis and Transform Methods参考文献 68被引用 7
一句话总结

本文提出了 $(f,\Gamma)$-散度,这是一种统一框架,介于 $f$-散度与积分概率度量(IPMs)之间,结合了二者的优势:既能处理非绝对连续测度(如 IPMs 所能),又能实现变分表示中的严格凹性(如 $f$-散度所具备)。该方法实现了对重尾分布和奇异分布的生成对抗网络(GAN)的稳定训练,在图像生成任务中表现优于梯度惩罚的Wasserstein GAN。

ABSTRACT

We develop a rigorous and general framework for constructing information-theoretic divergences that subsume both $f$-divergences and integral probability metrics (IPMs), such as the $1$-Wasserstein distance. We prove under which assumptions these divergences, hereafter referred to as $(f,Γ)$-divergences, provide a notion of `distance' between probability measures and show that they can be expressed as a two-stage mass-redistribution/mass-transport process. The $(f,Γ)$-divergences inherit features from IPMs, such as the ability to compare distributions which are not absolutely continuous, as well as from $f$-divergences, namely the strict concavity of their variational representations and the ability to control heavy-tailed distributions for particular choices of $f$. When combined, these features establish a divergence with improved properties for estimation, statistical learning, and uncertainty quantification applications. Using statistical learning as an example, we demonstrate their advantage in training generative adversarial networks (GANs) for heavy-tailed, not-absolutely continuous sample distributions. We also show improved performance and stability over gradient-penalized Wasserstein GAN in image generation.

研究动机与目标

  • 开发一种统一的散度框架,将 $f$-散度与积分概率度量(IPMs)相结合,以克服各自在统计学习中的局限性。
  • 实现对概率测度的鲁棒估计与比较,特别是在分布非绝对连续或具有重尾时。
  • 通过利用受约束函数空间的变分表示,提升生成建模中的训练稳定性和性能,尤其是针对 GAN。
  • 建立 $(f,\Gamma)$-散度的理论性质,包括严格凹性及两阶段质量重分配解释。
  • 在图像生成和分布建模任务中,通过实证验证其优于现有方法(如梯度惩罚的Wasserstein GAN)的优越性。

提出的方法

  • 通过变分公式提出 $(f,\Gamma)$-散度:$D_f^\Gamma(Q\|P) = \sup_{g \in \Gamma} \left\{ \mathbb{E}_Q[g] - \Lambda_f^P[g] \right\}$,其中 $\Lambda_f^P[g] = \inf_{\nu \in \mathbb{R}} \left\{ \nu + \mathbb{E}_P[f^*(g - \nu)] \right\}$。
  • 引入两阶段质量重分配过程:首先通过 $\nu$-优化选择最优偏移 $\nu$,然后在 $\Gamma$ 上优化 $g$,从而实现灵活且稳定的散度估计。
  • 利用凸函数 $f$ 的 Legendre 变换 $f^*$,且满足 $f(1) = 0$,以确保正确的散度结构与变分对偶性。
  • 采用有界可测函数 $\Gamma \subset \mathcal{M}_b(\Omega)$ 以保证问题的适定性,并在实践中可通过神经网络实现近似。
  • 通过分析二阶导数,推导出变分目标的严格凹性,从而确保优化过程的收敛性。
  • 将该框架应用于 GAN 训练,使用神经网络作为判别器,其中 $\Gamma$ 受限以强制实现利普希茨或反向利普希茨条件,提升稳定性。

实验结果

研究问题

  • RQ1能否构建一个统一的散度框架,使其兼具 IPMs 的鲁棒性与 $f$-散度的强变分性质?
  • RQ2在变分表示中引入偏移优化步骤($\nu$)如何影响散度的理论与实际性质?
  • RQ3$(f,\Gamma)$-散度是否能稳定 GAN 在重尾或奇异分布上的训练,而这些情况下标准 $f$-散度会失效?
  • RQ4函数空间 $\Gamma$ 在控制散度行为方面起什么作用,特别是在非绝对连续设定下?
  • RQ5$(f,\Gamma)$-散度的两阶段质量重分配解释如何揭示其几何与统计性质?

主要发现

  • 在 $f$ 与 $\Gamma$ 满足温和正则性条件时,$(f,\Gamma)$-散度框架为概率测度之间提供了一种距离概念,广义化了 $f$-散度与 IPMs。
  • 当测度 $P_0$ 下扰动 $\psi$ 的方差非零时,$(f,\Gamma)$-散度的变分表示在 $g$ 上是严格凹的,从而确保优化过程的稳定性。
  • 在 KL 散度情形下,目标函数的二阶导数为 $-\operatorname{Var}_{P_0}[\psi]$,证实了严格凹性,并在对偶形式中提供了收敛性保证。
  • 该框架能够有效训练 GAN,用于重尾和非绝对连续分布,而经典 $f$-GAN 在此类情形下即使散度有限也会失效。
  • 实证结果表明,在图像生成任务中,该方法在性能与稳定性上均优于梯度惩罚的Wasserstein GAN,收敛更快且生成样本质量更高。
  • 基于 $(f,\Gamma)$-散度的反向利普希茨 $\alpha$-GAN 变体,在样本保真度与训练稳定性方面优于标准 $f$-GAN 与 WGAN-GP,即使 $D_f(P_\theta\|Q) < \infty$ 时亦然。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。