Skip to main content
QUICK REVIEW

[论文解读] Debiaser Beware: Pitfalls of Centering Regularized Transport Maps

Aram-Alexandre Pooladian, Marco Cuturi|arXiv (Cornell University)|Feb 17, 2022
Groundwater flow and contamination studies被引用 6
一句话总结

本文研究了正则化最优传输映射去偏的统计权衡,表明尽管在正则化参数较小时去偏(通过Sinkhorn映射)能提升性能,但在正则化较强或样本量较小时反而会降低估计精度——这挑战了去偏在熵正则最优传输中普遍有益的广泛假设。

ABSTRACT

Estimating optimal transport (OT) maps (a.k.a. Monge maps) between two measures $P$ and $Q$ is a problem fraught with computational and statistical challenges. A promising approach lies in using the dual potential functions obtained when solving an entropy-regularized OT problem between samples $P_n$ and $Q_n$, which can be used to recover an approximately optimal map. The negentropy penalization in that scheme introduces, however, an estimation bias that grows with the regularization strength. A well-known remedy to debias such estimates, which has gained wide popularity among practitioners of regularized OT, is to center them, by subtracting auxiliary problems involving $P_n$ and itself, as well as $Q_n$ and itself. We do prove that, under favorable conditions on $P$ and $Q$, debiasing can yield better approximations to the Monge map. However, and perhaps surprisingly, we present a few cases in which debiasing is provably detrimental in a statistical sense, notably when the regularization strength is large or the number of samples is small. These claims are validated experimentally on synthetic and real datasets, and should reopen the debate on whether debiasing is needed when using entropic optimal transport.

研究动机与目标

  • 挑战去偏正则化最优传输映射始终提升估计精度的普遍假设。
  • 分析在不同正则化强度和样本量下,熵映射及其去偏变体(Sinkhorn映射)的统计特性。
  • 识别去偏在有限样本和高正则化情形下可证明性能下降的条件。
  • 提供理论与实证证据,表明去偏在正则化最优传输中不应被普遍应用。

提出的方法

  • 通过从每个测度与其自身之间的辅助最优传输问题中减去,定义去偏映射估计器。
  • 利用总体水平分析,在小正则化极限($\varepsilon \to 0$)下分析熵映射及其去偏变体的收敛速度。
  • 推导高斯到高斯传输映射情形下的闭式表达式,以实现熵映射与去偏映射的理论比较。
  • 建立理论界,表明当 $\varepsilon$ 较大或 $n$ 较小时,去偏映射的性能可能劣于熵映射。
  • 在合成数据集和真实世界数据集(包括单细胞基因组学数据)上进行数值实验,比较预测目标分布与真实目标分布之间的 $W_2$ 距离。
  • 使用POT库计算测试集上的无正则化 $W_2$ 距离,以验证有限样本效应。

实验结果

研究问题

  • RQ1在何种条件下,去偏会改善或降低正则化最优传输映射的统计性能?
  • RQ2正则化参数 $\varepsilon$ 如何影响熵映射与去偏Sinkhorn映射的相对性能?
  • RQ3在高维或低样本量情形下,去偏的有限样本效应是什么?
  • RQ4在高斯到高斯情形下,去偏映射是否在所有 $\varepsilon$ 值下均一致优于熵映射?
  • RQ5即使正则化参数较小,由于样本量限制,去偏是否仍可能导致更差的估计?

主要发现

  • 当正则化参数 $\varepsilon$ 较大或样本数 $n$ 较小时,去偏可能降低统计性能,这与普遍认知相反。
  • 在高斯到高斯情形下,当 $\varepsilon \to 0$ 时,去偏映射对Monge映射的逼近严格优于熵映射。
  • 在有限样本下,特别是在高维情形($d=15$)中,去偏映射表现出显著的有限样本效应,导致其性能劣于熵映射。
  • 在单细胞基因组学数据上的实证结果表明,随着 $\varepsilon$ 减小,熵映射在与目标分布的 $W_2$ 距离上优于去偏映射。
  • 当目标分布具有低方差分量时,熵映射与去偏映射之间的性能差距被放大,表明去偏在这些情形下会加剧偏差。
  • 去偏不应对所有情况一概而论;其优势取决于 $\varepsilon$ 和 $n$,且对正则化强度选择不佳或样本量较小的情形不具鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。