[论文解读] Universal Approximation for Log-concave Distributions using Well-conditioned Normalizing Flows
本文证明,当输入通过独立高斯分布进行填充时,任何对数凹分布均可被良好条件化的仿射耦合归一化流普遍近似,从而解决了归一化流表示理论中的一个关键空白。作者将仿射耦合与欠阻尼朗之万动力学及Hénon型映射联系起来,为高斯填充在训练稳定、良好条件化流时的实证成功提供了理论依据。
Normalizing flows are a widely used class of latent-variable generative models with a tractable likelihood. Affine-coupling (Dinh et al, 2014-16) models are a particularly common type of normalizing flows, for which the Jacobian of the latent-to-observable-variable transformation is triangular, allowing the likelihood to be computed in linear time. Despite the widespread usage of affine couplings, the special structure of the architecture makes understanding their representational power challenging. The question of universal approximation was only recently resolved by three parallel papers (Huang et al.,2020;Zhang et al.,2020;Koehler et al.,2020) -- who showed reasonably regular distributions can be approximated arbitrarily well using affine couplings -- albeit with networks with a nearly-singular Jacobian. As ill-conditioned Jacobians are an obstacle for likelihood-based training, the fundamental question remains: which distributions can be approximated using well-conditioned affine coupling flows? In this paper, we show that any log-concave distribution can be approximated using well-conditioned affine-coupling flows. In terms of proof techniques, we uncover and leverage deep connections between affine coupling architectures, underdamped Langevin dynamics (a stochastic differential equation often used to sample from Gibbs measures) and H\\'enon maps (a structured dynamical system that appears in the study of symplectic diffeomorphisms). Our results also inform the practice of training affine couplings: we approximate a padded version of the input distribution with iid Gaussians -- a strategy which Koehler et al.(2020) empirically observed to result in better-conditioned flows, but had hitherto no theoretical grounding. Our proof can thus be seen as providing theoretical evidence for the benefits of Gaussian padding when training normalizing flows.
研究动机与目标
- 解决良好条件化的仿射耦合归一化流是否能普遍近似行为良好的分布这一问题。
- 为高斯填充在训练稳定归一化流时的实证成功提供理论依据,相较于零填充或无填充。
- 确立对数凹分布——在实践中常见——可被良好条件化的流近似。
- 揭示仿射耦合架构、欠阻尼朗之万动力学与Hénon型动力系统之间的深层数学联系。
提出的方法
- 利用仿射耦合流与欠阻尼朗之万动力学之间的联系,后者是一种用于采样吉布斯测度的随机微分方程。
- 将流视为确定性动力系统的微扰,使用常微分方程流映射及通过变分方程进行的敏感性分析。
- 应用Grönwall不等式,以在小扰动下界定流映射中的误差传播,确保稳定性和良好条件性。
- 通过流映射的微扰展开,证明真实流与近似流之间的差异为$O(\epsilon^2)$,从而确保高精度。
- 引入一种填充输入分布,其中原始数据通过独立高斯分布增强,以实现良好条件化的雅可比行列式。
- 采用基于轨迹的流动力学分析,将流视为受控非线性性支配的常微分方程系统的时间演化变换。
实验结果
研究问题
- RQ1良好条件化的仿射耦合流能否普遍近似对数凹分布?
- RQ2为何高斯填充能提升归一化流的训练稳定性,是否存在理论依据?
- RQ3仿射耦合流、欠阻尼朗之万动力学与Hénon型映射之间存在何种数学联系?
- RQ4即使对于光滑、行为良好的分布,是否也能在避免病态雅可比行列式的情况下实现普遍近似?
主要发现
- 在$\mathbb{R}^d$中,任何对数凹分布均可通过在输入中添加独立高斯分布,被具有良好条件雅可比行列式的仿射耦合流任意精确地近似。
- 在紧致时间区间上,真实流与扰动流之间的近似误差在$C^r$拓扑下为$O(\epsilon^2)$,确保高精度。
- 无论目标分布的光滑性如何,雅可比行列式的条件数始终保持有界,解决了先前普遍近似结果中的一个关键局限。
- 理论分析证实,高斯填充相比零填充或无填充,能产生更良好条件化的流,解释了其在实证中的成功。
- 建立了仿射耦合流、欠阻尼朗之万动力学与Hénon型映射之间的深层联系,揭示了统一的动力系统视角。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。