Skip to main content
QUICK REVIEW

[论文解读] Singularity, Misspecification, and the Convergence Rate of EM

Raaz Dwivedi, Nhat Ho|arXiv (Cornell University)|Oct 1, 2018
Bayesian Methods and Mixture Models参考文献 33被引用 15
一句话总结

本文分析了在过度指定的高斯混合模型中,期望最大化(EM)算法的收敛行为,其中拟合的分量数量超过真实数量。研究揭示了一个明显的二元对立:在不平衡混合模型中,EM 算法以几何速度收敛至距离真实参数 $\mathcal{O}((d/n)^{1/2})$ 的点;而在平衡混合模型中——其中费雪信息矩阵变得奇异——收敛速度减缓至 $\mathcal{O}((d/n)^{1/4})$,与模型误设下的最大似然估计(MLE)的非渐近收敛速率一致。

ABSTRACT

A line of recent work has analyzed the behavior of the Expectation-Maximization (EM) algorithm in the well-specified setting, in which the population likelihood is locally strongly concave around its maximizing argument. Examples include suitably separated Gaussian mixture models and mixtures of linear regressions. We consider over-specified settings in which the number of fitted components is larger than the number of components in the true distribution. Such misspecified settings can lead to singularity in the Fisher information matrix, and moreover, the maximum likelihood estimator based on $n$ i.i.d. samples in $d$ dimensions can have a non-standard $\mathcal{O}((d/n)^{\frac{1}{4}})$ rate of convergence. Focusing on the simple setting of two-component mixtures fit to a $d$-dimensional Gaussian distribution, we study the behavior of the EM algorithm both when the mixture weights are different (unbalanced case), and are equal (balanced case). Our analysis reveals a sharp distinction between these two cases: in the former, the EM algorithm converges geometrically to a point at Euclidean distance of $\mathcal{O}((d/n)^{\frac{1}{2}})$ from the true parameter, whereas in the latter case, the convergence rate is exponentially slower, and the fixed point has a much lower $\mathcal{O}((d/n)^{\frac{1}{4}})$ accuracy. Analysis of this singular case requires the introduction of some novel techniques: in particular, we make use of a careful form of localization in the associated empirical process, and develop a recursive argument to progressively sharpen the statistical rate.

研究动机与目标

  • 理解在模型误设情况下,拟合过度指定混合模型时 EM 算法的计算与统计行为。
  • 研究费雪信息矩阵的奇异性如何影响高维设置下的收敛速率。
  • 刻画平衡与非平衡双分量高斯混合模型之间收敛行为的根本差异。
  • 在过度指定条件下,建立 EM 算法的非渐近收敛保证,尤其是在奇异情形下。
  • 开发新颖的分析技术,包括经验过程中的局部化方法与递归速率优化,以应对非标准的收敛情形。

提出的方法

  • 分析拟合至真实单分量高斯分布数据的双分量高斯混合模型中的总体 EM 算子。
  • 通过泰勒展开及正交旋转将问题转化为一维问题,推导总体 EM 算子 $\overline{M}(\theta)$ 的界。
  • 基于参数空间中的环形区域应用局部化论证,逐步优化收敛速率。
  • 利用总体 EM 的次几何收敛性,推导样本 EM 的非渐近收敛速率。
  • 运用集中不等式与经验过程理论,控制样本算子与总体算子之间的偏差。
  • 引入递归论证,将统计速率从 $\mathcal{O}((d/n)^{1/4})$ 锐化至最优的非渐近界。

实验结果

研究问题

  • RQ1当 EM 算法拟合双分量高斯混合模型至由单分量高斯分布生成的数据时,其行为如何?
  • RQ2混合平衡性(等重与非等重)对过度指定设置下 EM 收敛速率有何影响?
  • RQ3为何在平衡情况下收敛速率退化为 $\mathcal{O}((d/n)^{1/4})$?这与费雪信息矩阵的奇异性有何关联?
  • RQ4新颖的局部化与递归技术能否改善模型误设下 EM 的非渐近收敛分析?
  • RQ5在混合模型中,正确设定与过度指定且奇异的情形之间,收敛行为的根本差异是什么?

主要发现

  • 在非平衡情况下,EM 算法以几何速度收敛至距离真实参数 $\mathcal{O}((d/n)^{1/2})$ 的欧几里得距离内。
  • 在平衡情况下,EM 的收敛速度显著变慢,仅能达到 $\mathcal{O}((d/n)^{1/4})$ 的精度,原因在于费雪信息矩阵的奇异性。
  • 总体 EM 算子满足 $\|\overline{M}(\theta)\|_2 \in \left[\|\theta\|_2(1 - 3\|\theta\|_2^2), \|\theta\|_2(1 - 2\|\theta\|_2^2)\right]$,当 $\|\theta\|_2 \leq 1/2$ 时,表明在原点附近收缩缓慢。
  • 在适当的样本量 $n \geq c'_1 d \log(\log(1/\epsilon)/\delta)$ 下,样本 EM 算法以高概率实现 $\mathcal{O}((d/n)^{1/4 - \epsilon})$ 的收敛速率。
  • 分析表明,$\mathcal{O}((d/n)^{1/4})$ 的速率是紧致的,且与已知的过度指定下 MLE 的非渐近收敛速率一致。
  • 提出一种新颖的递归局部化论证,可逐步优化统计速率,从而克服奇异费雪信息矩阵带来的挑战。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。