[论文解读] Estimation and inference for the Wasserstein distance between mixing measures in topic models
本文首次为主题模型中混合测度间Wasserstein距离的估计提供了极小极大下界,通过公理化论证确立了其作为标准度量的地位。此外,本文进一步开发了完全基于数据的推断工具,实现了在高维、可能稀疏的主题模型中Wasserstein距离的渐近有效置信区间。
The Wasserstein distance between mixing measures has come to occupy a central place in the statistical analysis of mixture models. This work proposes a new canonical interpretation of this distance and provides tools to perform inference on the Wasserstein distance between mixing measures in topic models. We consider the general setting of an identifiable mixture model consisting of mixtures of distributions from a set $\mathcal{A}$ equipped with an arbitrary metric $d$, and show that the Wasserstein distance between mixing measures is uniquely characterized as the most discriminative convex extension of the metric $d$ to the set of mixtures of elements of $\mathcal{A}$. The Wasserstein distance between mixing measures has been widely used in the study of such models, but without axiomatic justification. Our results establish this metric to be a canonical choice. Specializing our results to topic models, we consider estimation and inference of this distance. Though upper bounds for its estimation have been recently established elsewhere, we prove the first minimax lower bounds for the estimation of the Wasserstein distance in topic models. We also establish fully data-driven inferential tools for the Wasserstein distance in the topic model context. Our results apply to potentially sparse mixtures of high-dimensional discrete probability distributions. These results allow us to obtain the first asymptotically valid confidence intervals for the Wasserstein distance in topic models.
研究动机与目标
- 为在可识别的混合模型中将混合测度间的Wasserstein距离作为标准度量使用提供公理化基础。
- 解决在同时估计成分和权重时,主题模型中混合测度间Wasserstein距离的极小极大下界这一开放问题。
- 为具有高维、稀疏离散分布的主题模型中的Wasserstein距离,开发完全基于数据的推断程序,包括渐近有效的置信区间。
- 通过建立该估计问题的首个极小极大下界,弥合现有上界与最优估计速率之间的差距。
- 将理论工具扩展至混合成分是离散且可能低秩的稀疏高维设置中。
提出的方法
- 证明混合测度间的Wasserstein距离是集合 $ \mathcal{A} $ 上基度量 $ d $ 的唯一最具区分力的凸扩展,从而提供公理化依据。
- 使用约化方案,通过将 $ W(\widehat{\boldsymbol{\alpha}}, \boldsymbol{\alpha}; d) $ 的估计与 $ \widehat{\alpha} - \alpha $ 和 $ \widehat{A} - A $ 的 $ \ell_1 $-范数估计关联,推导极小极大下界。
- 利用已知的在稀疏性约束下对高维稀疏向量和矩阵分量进行 $ \ell_1 $-估计的结果。
- 将约化应用于两个不同的子问题:一个聚焦于固定成分时的 $ \|\widehat{\alpha} - \alpha\|_1 $,另一个聚焦于固定权重时的 $ \|\widehat{A} - A\|_1 $。
- 在正则性条件下,利用推导出的极小极大速率和估计量的渐近正态性,构建基于数据的推断程序。
- 利用满足 $ \min_{k \neq k'} d(A_k, A_{k'}) \geq c_d \overline{\kappa}_\tau \|A_k - A_{k'}\|_1 $ 的度量 $ d $,以控制成分之间的分离性。
实验结果
研究问题
- RQ1混合测度间的Wasserstein距离是否是比较混合模型的合理选择?如果是,其公理化依据是什么?
- RQ2当同时估计成分分布和混合权重时,主题模型中混合测度间Wasserstein距离的极小极大下界是什么?
- RQ3能否为主题模型中混合测度间的Wasserstein距离构建完全基于数据的置信区间?
- RQ4在高维离散设置中,估计速率如何随模型复杂度(如稀疏性 $ \tau $、样本量 $ n $ 和维度 $ pK $)变化?
- RQ5在稀疏高维主题模型中,估计混合测度间Wasserstein距离的最优收敛速率是什么?
主要发现
- 混合测度间的Wasserstein距离被唯一表征为基度量 $ d $ 的最具区分力的凸扩展,为其使用提供了公理化基础。
- 本文首次建立了主题模型中估计Wasserstein距离的极小极大下界:$ \inf_{\widehat{\boldsymbol{\alpha}}}\sup_{\alpha \in \Theta_{\alpha}(\tau), A \in \Theta_A} \mathbb{E}[W(\widehat{\boldsymbol{\alpha}}, \boldsymbol{\alpha}; d)] \gtrsim c_d \overline{\kappa}_\tau \left( \sqrt{\tau/N} + \frac{1}{\overline{\kappa}_\tau} \sqrt{pK/(nN)} \right) $。
- 该下界表明,近似极小极大最优估计是可实现的,与文献中近期的上界结果相辅相成。
- 作者构建了首个完全基于数据的推断工具,使得在主题模型中Wasserstein距离的渐近有效置信区间成为可能。
- 结果在一般条件下成立,包括可能稀疏的高维离散分布混合,其中 $ 1 < \tau \leq cN $ 且 $ pK \leq c'(nN) $。
- 理论框架适用于满足 $ \min_{k \neq k'} d(A_k, A_{k'}) \geq c_d \overline{\kappa}_\tau \|A_k - A_{k'}\|_1 $ 的广泛类度量 $ d $,从而确保成分之间具有有意义的分离性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。