[Paper Review] Cumulant-free closed-form formulas for some common (dis)similarities between densities of an exponential family
This paper presents cumulant-free closed-form formulas for common (dis)similarities—such as Kullback-Leibler, Bhattacharyya, Hellinger, and α-divergences—between densities in exponential families. By leveraging quasi-arithmetic means and canonical parameterization, the method bypasses explicit computation of the cumulant function, enabling efficient implementation via legacy APIs using only partial density factorization and inverse parameter mapping.
It is well-known that the Bhattacharyya, Hellinger, Kullback-Leibler, $α$-divergences, and Jeffreys' divergences between densities belonging to a same exponential family have generic closed-form formulas relying on the strictly convex and real-analytic cumulant function characterizing the exponential family. In this work, we report (dis)similarity formulas which bypass the explicit use of the cumulant function and highlight the role of quasi-arithmetic means and their multivariate mean operator extensions. In practice, these cumulant-free formulas are handy when implementing these (dis)similarities using legacy Application Programming Interfaces (APIs) since our method requires only to partially factorize the densities canonically of the considered exponential family.
Motivation & Objective
- To derive closed-form expressions for key statistical (dis)similarities between densities in exponential families without explicit use of the cumulant function.
- To enable practical computation of divergences like Kullback-Leibler, Bhattacharyya, and Hellinger using only standard software APIs and partial density factorization.
- To demonstrate that these (dis)similarities can be computed via generalized weighted quasi-arithmetic means derived from the natural parameter mapping.
- To provide alternative computational pathways for divergences when the cumulant function is intractable or unavailable.
Proposed method
- The method uses the canonical parameterization of exponential family densities to express (dis)similarities in terms of the inverse of the natural parameter mapping θ(λ) and its associated quasi-arithmetic mean operator.
- It formulates the Bhattacharyya coefficient and related divergences using a reference point ω in the support X, allowing the use of p(ω;λ) as a basis for computation.
- For Kullback-Leibler divergence, the approach uses a limit-based expression involving α-skew Bhattacharyya distances or a first-order approximation of the weighted mean to avoid cumulant evaluation.
- It also employs a Legendre-Fenchel duality formulation that expresses KL divergence in terms of entropy, moments, and log-density ratios, avoiding the cumulant function.
- The method further allows expressing KL divergence as a weighted sum of log-density ratios over a set of s ≤ D+1 sample points ω_i that match the first-order moment of the source distribution.
- Implementation relies on standard parametric density libraries and avoids symbolic integration by using pointwise evaluations and inverse parameter functions.
Experimental results
Research questions
- RQ1Can common (dis)similarities between exponential family densities be expressed without explicitly computing the cumulant function?
- RQ2How can quasi-arithmetic means be used to construct closed-form formulas for divergences like Bhattacharyya and Hellinger?
- RQ3What alternative computational pathways exist for Kullback-Leibler divergence when the cumulant function is intractable?
- RQ4Can a finite set of sample points ω_i be used to exactly represent the KL divergence via log-density ratios?
Key findings
- The Bhattacharyya coefficient and related divergences (Hellinger, α-divergences) can be computed using any point ω in the support X, with the formula depending only on p(ω;λ) and the generalized quasi-arithmetic mean of the parameters.
- The Kullback-Leibler divergence can be expressed as a limit of α-skew Bhattacharyya distances, avoiding the cumulant function through symbolic or numerical approximation.
- For univariate and multivariate Gaussian families, the required sample points ω_i for the KL divergence sum formula can be explicitly constructed to match the first-order moment.
- The KL divergence can be rewritten as a sum of log-density ratios over s ≤ D+1 points, where D is the dimension of the parameter space, enabling exact computation without cumulant evaluation.
- The Jeffreys’ divergence is shown to equal the inner product of the difference in natural parameters and the difference in sufficient statistics' expectations, a form independent of the cumulant function.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.