Skip to main content
QUICK REVIEW

[論文レビュー] Singularity, Misspecification, and the Convergence Rate of EM

Raaz Dwivedi, Nhat Ho|arXiv (Cornell University)|Oct 1, 2018
Bayesian Methods and Mixture Models参考文献 33被引用数 15
ひとこと要約

本稿は、真の成分数より多く指定されたガウス・ミックスチャネル・モデルにおける期待最大化(EM)アルゴリズムの収束挙動を分析する。過剰適合された状況では、混合成分の重みが不均衡な場合、EMは幾何的収束を示し、真のパラメータからのユークリッド距離が $\mathcal{O}((d/n)^{1/2})$ の範囲内に収束する。一方、混合成分の重みが均衡している場合、フィッシャー情報行列が特異化するため、収束速度は $\mathcal{O}((d/n)^{1/4})$ に低下し、モデル不適合下でのMLEの非漸近的収束速度と一致する。

ABSTRACT

A line of recent work has analyzed the behavior of the Expectation-Maximization (EM) algorithm in the well-specified setting, in which the population likelihood is locally strongly concave around its maximizing argument. Examples include suitably separated Gaussian mixture models and mixtures of linear regressions. We consider over-specified settings in which the number of fitted components is larger than the number of components in the true distribution. Such misspecified settings can lead to singularity in the Fisher information matrix, and moreover, the maximum likelihood estimator based on $n$ i.i.d. samples in $d$ dimensions can have a non-standard $\mathcal{O}((d/n)^{\frac{1}{4}})$ rate of convergence. Focusing on the simple setting of two-component mixtures fit to a $d$-dimensional Gaussian distribution, we study the behavior of the EM algorithm both when the mixture weights are different (unbalanced case), and are equal (balanced case). Our analysis reveals a sharp distinction between these two cases: in the former, the EM algorithm converges geometrically to a point at Euclidean distance of $\mathcal{O}((d/n)^{\frac{1}{2}})$ from the true parameter, whereas in the latter case, the convergence rate is exponentially slower, and the fixed point has a much lower $\mathcal{O}((d/n)^{\frac{1}{4}})$ accuracy. Analysis of this singular case requires the introduction of some novel techniques: in particular, we make use of a careful form of localization in the associated empirical process, and develop a recursive argument to progressively sharpen the statistical rate.

研究の動機と目的

  • 過剰適合混合モデルをフィッティングする際のEMアルゴリズムの計算的・統計的挙動を理解すること。特に、モデル不適合が生じる状況での挙動を対象とする。
  • 高次元設定におけるフィッシャー情報行列の特異性が収束速度に与える影響を調査すること。
  • 2成分ガウス混合モデルにおけるバランス(等重み vs. 異重み)の有無による収束挙動の根本的差を特定すること。
  • 特に特異な状況における過剰適合下でのEMの非漸近的収束保証を確立すること。
  • 非標準的な収束状態に対処するための新規解析的手法、特に経験過程における局所化と再帰的レート鋭化を構築すること。

提案手法

  • 真の1成分ガウス分布から生成されたデータに2成分ガウス混合モデルをフィッティングした場合の母集団EM作用素を分析する。
  • テイラー展開と正規直交回転を用いて1次元問題に変換することで、母集団EM作用素 $\overline{M}(\theta)$ の境界を導出する。
  • パrameter空間におけるアニュラスに基づく局所化の議論を用い、段階的に収束速度を改善する。
  • 母集団EMの非幾何的収束を用いて、標本EMの非漸近的収束速度を導出する。
  • 集中不等式と経験過程理論を用い、経験的作用素と母集団作用素の乖離を制御する。
  • 再帰的議論を導入し、統計的レートを $\mathcal{O}((d/n)^{1/4})$ から最適な非漸近的境界へと鋭くする。

実験結果

リサーチクエスチョン

  • RQ11成分ガウス分布から生成されたデータに2成分ガウス混合モデルをフィッティングするEMアルゴリズムは、どのように振る舞うか?
  • RQ2混合成分の重みのバランス(等重み vs. 異重み)は、過剰適合状況下でのEMの収束速度にどのような影響を与えるか?
  • RQ3バランス状態ではなぜ収束速度が $\mathcal{O}((d/n)^{1/4})$ に低下するのか?これはフィッシャー情報行列の特異性とどのように関係しているか?
  • RQ4新規の局所化および再帰的技法は、モデル不適合下でのEMの非漸近的収束解析をどのように改善できるか?
  • RQ5混合モデルにおいて、正しく指定された場合と過剰適合・特異な状況との間で、収束挙動の根本的差は何か?

主な発見

  • 不均衡状況では、EMは幾何的収束を示し、真のパラメータからのユークリッド距離が $\mathcal{O}((d/n)^{1/2})$ の範囲内に収束する。
  • 均衡状況では、収束速度が指数的に遅くなり、フィッシャー情報行列の特異性のため、$\mathcal{O}((d/n)^{1/4})$ の精度しか達成できない。
  • 母集団EM作用素は、$\|\theta\|_2 \leq 1/2$ の範囲で $\|\overline{M}(\theta)\|_2 \in \left[\|\theta\|_2(1 - 3\|\theta\|_2^2), \|\theta\|_2(1 - 2\|\theta\|_2^2)\right]$ を満たし、原点付近での収縮が遅いことが示唆される。
  • 適切な標本サイズ $n \geq c'_1 d \log(\log(1/\epsilon)/\delta)$ の下で、標本EMアルゴリズムは高確率で $\mathcal{O}((d/n)^{1/4 - \epsilon})$ のレートを達成する。
  • 解析により、$\mathcal{O}((d/n)^{1/4})$ のレートがタイトであり、過剰適合下でのMLEの既知の非漸近的収束速度と一致することが示された。
  • 特異なフィッシャー情報行列が引き起こす課題に対処するため、段階的に統計的レートを鋭くする新規の再帰的局所化議論が開発された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。