Skip to main content
QUICK REVIEW

[论文解读] Data-driven regularization of Wasserstein barycenters with an application to multivariate density registration

Jérémie Bigot, Elsa Cazelles|arXiv (Cornell University)|Apr 24, 2018
Point processes and geometric inequalities参考文献 28被引用 7
一句话总结

该论文提出了一种基于数据的正则化框架,用于Wasserstein中位数,以同时对来自噪声且错位的密度数据的多变量点云进行对齐与平滑。通过将Goldenshluger-Lepski原则应用于熵正则化(Sinkhorn中位数)和函数惩罚方法中的正则化参数选择,该方法在密度配准中提高了鲁棒性和准确性,其有效性在模拟的高斯混合模型和真实的流式细胞术数据上得到验证。

ABSTRACT

We present a framework to simultaneously align and smooth data in the form of multiple point clouds sampled from unknown densities with support in a d-dimensional Euclidean space. This work is motivated by applications in bioinformatics where researchers aim to automatically homogenize large datasets to compare and analyze characteristics within a same cell population. Inconveniently, the information acquired is most certainly noisy due to mis-alignment caused by technical variations of the environment. To overcome this problem, we propose to register multiple point clouds by using the notion of regularized barycenters (or Fr\\'{e}chet mean) of a set of probability measures with respect to the Wasserstein metric. A first approach consists in penalizing a Wasserstein barycenter with a convex functional as recently proposed in Bigot and al. (2018). A second strategy is to transform the Wasserstein metric itself into an entropy regularized transportation cost between probability measures as introduced in Cuturi (2013). The main contribution of this work is to propose data-driven choices for the regularization parameters involved in each approach using the Goldenshluger-Lepski's principle. Simulated data sampled from Gaussian mixtures are used to illustrate each method, and an application to the analysis of flow cytometry data is finally proposed. This way of choosing of the regularization parameter for the Sinkhorn barycenter is also analyzed through the prism of an oracle inequality that relates the error made by such data-driven estimators to the one of an ideal estimator.

研究动机与目标

  • 解决由于数据采集过程中的技术性错位导致的多变量密度数据相位可变性,特别是在高维生物数据集中。
  • 开发一个统一框架,以同时配准并平滑从 $\mathbb{R}^d$ 中未知概率密度采样的多个点云。
  • 通过使用Wasserstein中位数作为几何均值,克服标准欧几里得平均因错位而无法保持形状的局限性。
  • 引入基于Goldenshluger-Lepski原则的数据驱动正则化参数选择方法,以在不依赖底层密度结构先验知识的情况下优化性能。
  • 在合成数据(高斯混合模型)和真实世界流式细胞术数据上验证该方法,以展示其在生物信息学中的实际应用价值。

提出的方法

  • 将Wasserstein中位数表述为在 $W_2$ 距离下的概率测度空间中的Fréchet均值,从而实现对错位点云的对齐与平滑。
  • 应用两种正则化策略:(1) 通过Sinkhorn散度实现熵正则化,使Wasserstein距离在计算上可行;(2) 通过函数惩罚(如TV、Tikhonov)实现平滑性约束。
  • 使用Goldenshluger-Lepski原则以数据驱动方式选择最优正则化参数 $\varepsilon$ 和 $\gamma$,以最小化估计风险。
  • 通过Sinkhorn算法实现对偶优化,并采用对偶变量重参数化($\psi_i = \phi_i + K^T\phi_0/n$)以稳定收敛并降低计算复杂度。
  • 采用次梯度下降法(光滑惩罚使用L-BFGS,非光滑惩罚使用FISTA),并以 $K=10$ 个最近邻计算运输映射的稳定次梯度。
  • 利用最优对偶变量通过前推公式 $f_i^k = \sum_j \nu_i^j S_i^{jk}$ 重构中位数,其中 $S_i$ 是基于最近邻构造的随机运输矩阵。

实验结果

研究问题

  • RQ1如何同时对齐并平滑多个错位的多变量点云,以恢复一个具有代表性且形状保持的中位数?
  • RQ2在保持数据保真度与结果密度平滑性之间取得平衡的前提下,正则化Wasserstein中位数的最优方式是什么?
  • RQ3Goldenshluger-Lepski原则能否在熵正则化与函数正则化背景下有效应用于正则化参数的数据驱动选择?
  • RQ4与已知最优参数的oracle估计器相比,数据驱动正则化方法的性能如何?
  • RQ5所提出的方法在真实世界生物数据集(如流式细胞术)中的配准精度提升程度如何?

主要发现

  • 通过Goldenshluger-Lepski原则实现的数据驱动正则化参数选择,使得中位数估计性能可与oracle估计器相媲美,其误差与理想估计器之间的关系由一个oracle不等式所描述。
  • 在模拟的高斯混合模型上,使用数据驱动 $\varepsilon=1.6$ 的Sinkhorn中位数成功恢复了真实底层密度形状,避免了欧几里得均值中出现的虚假模式。
  • 该方法有效减少了多变量密度数据中的相位可变性,如在不同受试者间对双峰和多峰分布的对齐效果所示。
  • 在次梯度计算中使用 $K=10$ 个最近邻显著提高了数值稳定性,相较于单个最近邻选择方法。
  • 该框架在真实流式细胞术数据上实现了精确的配准与平滑,从而实现了跨样本细胞群的可靠比较。
  • 理论分析表明,在较弱正则性假设下,所提出的估计器可达到最优收敛速率,其误差被限制为oracle风险的倍数。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。