[Paper Review] On the Universality of Volume-Preserving and Coupling-Based Normalizing Flows
This paper establishes a new theoretical framework proving that affine coupling-based normalizing flows are universal approximators for probability distributions, overcoming prior limitations tied to ill-conditioned networks and volume-preserving flows. The authors introduce a constructive, layer-by-layer training approach that ensures convergence to target distributions, demonstrating that expressive coupling functions enable efficient, practical universality without requiring pathological network behaviors.
We present a novel theoretical framework for understanding the expressive power of normalizing flows. Despite their prevalence in scientific applications, a comprehensive understanding of flows remains elusive due to their restricted architectures. Existing theorems fall short as they require the use of arbitrarily ill-conditioned neural networks, limiting practical applicability. We propose a distributional universality theorem for well-conditioned coupling-based normalizing flows such as RealNVP. In addition, we show that volume-preserving normalizing flows are not universal, what distribution they learn instead, and how to fix their expressivity. Our results support the general wisdom that affine and related couplings are expressive and in general outperform volume-preserving flows, bridging a gap between empirical results and theoretical understanding.
Motivation & Objective
- To address the gap between empirical success and theoretical understanding of coupling-based normalizing flows.
- To identify and resolve fundamental limitations in existing universality theorems that rely on ill-conditioned neural networks.
- To demonstrate that volume-preserving flows are not universal approximators under KL divergence, a key practical loss measure.
- To provide a constructive, practical universality proof for affine coupling flows using sequential layer training.
- To guide the design of more expressive coupling functions by clarifying their role in achieving efficient distribution approximation.
Proposed method
- Proposes a constructive universality proof for coupling-based normalizing flows by training layers sequentially, ensuring convergence to the target distribution.
- Uses affine coupling blocks that rotate the latent distribution and then transform active dimensions to zero mean and unit variance via learned scale and shift parameters.
- Employs cubic spline interpolation to estimate conditional means and standard deviations from binned data, enabling smooth parameterization of coupling functions.
- Applies a step-size constraint via convex combination with the identity map to stabilize training and reduce artifacts from finite data.
- Uses iterative resampling of training data to prevent overfitting during sequential layer optimization.
- Compares volume-preserving flows (constant Jacobian determinant) with variable-Jacobian flows to isolate the impact of volume preservation on expressivity.

Experimental results
Research questions
- RQ1Can coupling-based normalizing flows achieve universal approximation of probability distributions without relying on ill-conditioned neural networks?
- RQ2Is the volume-preserving property inherent in existing universality proofs for coupling flows, and does it limit practical applicability?
- RQ3Can a constructive, layer-by-layer training procedure ensure convergence to a target distribution in practice?
- RQ4How does the expressivity of coupling functions affect the number of layers required for good distribution approximation?
- RQ5What is the role of Jacobian determinant variation in the expressive power of normalizing flows?
Key findings
- Volume-preserving normalizing flows are not universal approximators under KL divergence, invalidating their use in practical distribution learning despite theoretical claims.
- Existing universality proofs for coupling flows implicitly construct volume-preserving transformations, which limits their practical relevance due to restricted expressivity.
- The proposed layer-by-layer training method constructs a normalizing flow that converges to the target distribution, as demonstrated on a ring-shaped data distribution with 100 layers.
- The method achieves accurate density estimation and sampling by iteratively reducing the loss through rotation and affine coupling, with the final latent distribution converging to a standard normal.
- The flow with variable Jacobian determinant outperforms volume-preserving counterparts, confirming that non-volume-preserving transformations are essential for expressive power.
- Theoretical and empirical results validate that affine coupling blocks are a strong foundation for normalizing flows, especially when coupling functions are sufficiently expressive.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.