[Paper Review] $(f,Γ)$-Divergences: Interpolating between $f$-Divergences and Integral Probability Metrics
This paper introduces $(f,\Gamma)$-divergences, a unified framework that interpolates between $f$-divergences and integral probability metrics (IPMs), combining their strengths: the ability to handle non-absolutely continuous measures (like IPMs) and strict concavity in variational representations (like $f$-divergences). The method enables stable training of generative adversarial networks (GANs) for heavy-tailed and singular distributions, outperforming gradient-penalized Wasserstein GANs in image generation tasks.
We develop a rigorous and general framework for constructing information-theoretic divergences that subsume both $f$-divergences and integral probability metrics (IPMs), such as the $1$-Wasserstein distance. We prove under which assumptions these divergences, hereafter referred to as $(f,Γ)$-divergences, provide a notion of `distance' between probability measures and show that they can be expressed as a two-stage mass-redistribution/mass-transport process. The $(f,Γ)$-divergences inherit features from IPMs, such as the ability to compare distributions which are not absolutely continuous, as well as from $f$-divergences, namely the strict concavity of their variational representations and the ability to control heavy-tailed distributions for particular choices of $f$. When combined, these features establish a divergence with improved properties for estimation, statistical learning, and uncertainty quantification applications. Using statistical learning as an example, we demonstrate their advantage in training generative adversarial networks (GANs) for heavy-tailed, not-absolutely continuous sample distributions. We also show improved performance and stability over gradient-penalized Wasserstein GAN in image generation.
Motivation & Objective
- To develop a general framework for divergences that unify $f$-divergences and integral probability metrics (IPMs), overcoming limitations of each in statistical learning.
- To enable robust estimation and comparison of probability measures, especially when distributions are not absolutely continuous or have heavy tails.
- To improve training stability and performance in generative modeling, particularly for GANs, by leveraging variational representations with constrained function spaces.
- To establish theoretical properties of $(f,\Gamma)$-divergences, including strict concavity and two-stage mass redistribution interpretation.
- To demonstrate empirical superiority over existing methods, such as gradient-penalized Wasserstein GANs, in image generation and distributional modeling.
Proposed method
- Proposes $(f,\Gamma)$-divergences via a variational formulation: $D_f^\Gamma(Q\|P) = \sup_{g \in \Gamma} \left\{ \mathbb{E}_Q[g] - \Lambda_f^P[g] \right\}$, where $\Lambda_f^P[g] = \inf_{\nu \in \mathbb{R}} \left\{ \nu + \mathbb{E}_P[f^*(g - \nu)] \right\}$.
- Introduces a two-stage mass-redistribution process: first, optimal shift $\nu$ is chosen via $\nu$-optimization, then $g$ is optimized over $\Gamma$, enabling flexible and stable divergence estimation.
- Uses the Legendre transform $f^*$ of a convex function $f$ with $f(1) = 0$ to ensure proper divergence structure and variational duality.
- Employs bounded measurable functions $\Gamma \subset \mathcal{M}_b(\Omega)$ to ensure well-posedness and enable approximation via neural networks in practice.
- Derives strict concavity of the variational objective by analyzing second-order derivatives, ensuring convergence in optimization.
- Applies the framework to train GANs using neural networks as discriminators, with $\Gamma$ constrained to enforce Lipschitz or reverse-Lipschitz conditions for stability.
Experimental results
Research questions
- RQ1Can a unified divergence framework be constructed that inherits the robustness of IPMs and the strong variational properties of $f$-divergences?
- RQ2How does the inclusion of a shift-optimization step ($\nu$) in the variational representation affect the theoretical and practical properties of the divergence?
- RQ3Can $(f,\Gamma)$-divergences stabilize GAN training for heavy-tailed or singular distributions where standard $f$-divergences fail?
- RQ4What is the role of the function space $\Gamma$ in controlling the behavior of the divergence, especially in non-absolute continuous settings?
- RQ5How does the two-stage mass-redistribution interpretation of $(f,\Gamma)$-divergences inform its geometric and statistical properties?
Key findings
- The $(f,\Gamma)$-divergence framework provides a notion of distance between probability measures under mild regularity conditions on $f$ and $\Gamma$, generalizing both $f$-divergences and IPMs.
- The variational representation of $(f,\Gamma)$-divergences is strictly concave in $g$ when the variance of the perturbation $\psi$ under the measure $P_0$ is non-zero, ensuring stable optimization.
- For the KL divergence case, the second derivative of the objective is $-\operatorname{Var}_{P_0}[\psi]$, confirming strict concavity and convergence guarantees in the dual formulation.
- The framework enables effective training of GANs on heavy-tailed and non-absolutely continuous distributions, where classical $f$-GANs fail even when the divergence is finite.
- Empirical results show improved performance and stability over gradient-penalized Wasserstein GANs in image generation tasks, with faster convergence and better sample quality.
- The reverse Lipschitz $\alpha$-GAN variant based on $(f,\Gamma)$-divergences achieves better sample fidelity and training stability than standard $f$-GANs and WGAN-GP, even when $D_f(P_\theta\|Q) < \infty$.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.