Skip to main content
QUICK REVIEW

[論文レビュー] Challenges with EM in application to weakly identifiable mixture models.

Raaz Dwivedi, Nhat Ho|arXiv (Cornell University)|Feb 1, 2019
Bayesian Methods and Mixture Models参考文献 22被引用数 6
ひとこと要約

本稿は、弱識別可能な位置尺度混合モデルにおいて最尤推定が部分的誤差率を示す場合に、期待最大化(EM)アルゴリズムの収束が遅い理由を調査する。解析的にEM収束速度を特定し、1次元の場合に$n^{3/4}$ステップ、多次元設定では$(n/d)^{1/2}$ステップを要し、位置パラメータの推定誤差は$n^{-1/8}$、スケールパラメータは$n^{-1/4}$のオーダーであることを示す。

ABSTRACT

We study a class of weakly identifiable location-scale mixture models for which the maximum likelihood estimates based on $n$ i.i.d. samples are known to have lower accuracy than the classical $n^{- \frac{1}{2}}$ error. We investigate whether the Expectation-Maximization (EM) algorithm also converges slowly for these models. We first demonstrate via simulation studies a broad range of over-specified mixture models for which the EM algorithm converges very slowly, both in one and higher dimensions. We provide a complete analytical characterization of this behavior for fitting data generated from a multivariate standard normal distribution using two-component Gaussian mixture with varying location and scale parameters. Our results reveal distinct regimes in the convergence behavior of EM as a function of the dimension $d$. In the multivariate setting ($d \geq 2$), when the covariance matrix is constrained to a multiple of the identity matrix, the EM algorithm converges in order $(n/d)^{\frac{1}{2}}$ steps and returns estimates that are at a Euclidean distance of order ${(n/d)^{-\frac{1}{4}}}$ and ${ (n d)^{- \frac{1}{2}}}$ from the true location and scale parameter respectively. On the other hand, in the univariate setting ($d = 1$), the EM algorithm converges in order $n^{\frac{3}{4} }$ steps and returns estimates that are at a Euclidean distance of order ${ n^{- \frac{1}{8}}}$ and ${ n^{-\frac{1} {4}}}$ from the true location and scale parameter respectively. Establishing the slow rates in the univariate setting requires a novel localization argument with two stages, with each stage involving an epoch-based argument applied to a different surrogate EM operator at the population level. We also show multivariate ($d \geq 2$) examples, involving more general covariance matrices, that exhibit the same slow rates as the univariate case.

研究の動機と目的

  • 最大尤度推定が$n^{-1/2}$未満の精度を示す弱識別可能な混合モデルにおいて、EMがなぜ遅く収束するのかを理解すること。
  • パラメータが変化する1次元および多次元正規位置尺度混合モデルにおけるEM収束行動を分析すること。
  • 過剰に指定されたモデル下でのEMの収束速度および推定誤差を特徴づけること、特に真のデータ生成分布が標準正規分布である場合に注目すること。
  • 1次元ケースにおける母集団レベルのEM作用素に対して、段階的局在化アプローチとエポックベース解析を用いた新規な2段階法を確立すること。

提案手法

  • 過剰に指定された2成分正規混合モデルを1次元および高次元でシミュレーションし、EM収束の遅さを実証する。
  • 恒等行列を共分散に制約した2成分正規混合モデルに対して、多次元正規データの解析的特徴づけを実施する。
  • 母集団レベルでの異なる補助EM作用素のエポックベース解析を用いて、1次元ケースを扱うための2段階局在化アプローチを開発する。
  • 一般共分散行列を有する多次元設定への分析を拡張し、1次元ケースと同様の遅い収束速度が得られることを示す。
  • 漸近的解析を用いて、標本サイズ$n$と次元$d$を増加させる条件下で収束速度と推定誤差を導出する。$n$および$d$への明示的依存関係を示す。

実験結果

リサーチクエスチョン

  • RQ1弱識別可能な混合モデルにおいて、推定精度が$n^{-1/2}$未満である場合、EMアルゴリズムの収束速度はどのように振る舞うか?
  • RQ22成分正規混合モデルにおけるEMの収束速度は、1次元と多次元設定でどのように異なるか?
  • RQ3EMが1次元ケースでなぜ遅く収束するのか、そしてこれは新規な解析的枠組みによって説明可能か?
  • RQ4多次元モデルにおいて共分散行列が恒等行列の倍数に制限されない場合、遅い収束速度は継続するか?
  • RQ5一般共分散構造を有する多次元モデルでも、同様の遅い収束速度が観察されるか?

主な発見

  • 1次元設定($d = 1$)では、EMアルゴリズムは$n^{3/4}$ステップで収束し、位置パラメータの推定誤差は$n^{-1/8}$、スケールパラメータは$n^{-1/4}$のオーダーである。
  • 多次元設定($d \to \text{dim} \to \text{infty}$)では、共分散が恒等行列の倍数に制限される場合、EMは$(n/d)^{1/2}$ステップで収束し、位置パラメータの推定誤差は$(n/d)^{-1/4}$、スケールパラメータは$(n d)^{-1/2}$のオーダーである。
  • 1次元ケースにおける遅い収束は、母集団レベルでの補助EM作用素のエポックベース解析を含む新規な2段階局在化アプローチによって説明可能である。
  • 一般共分散行列を有する多次元モデルでも、1次元ケースと同様の遅い収束速度が観察され、この現象が恒等行列構造に限定されないことが示された。
  • 結果として、弱識別可能なモデルでは、モデルが過剰に指定されていても、EMは古典的な$n^{-1/2}$収束速度に到達しないことが明らかになった。
  • 本研究は、EMの収束行動が、パrameter空間の幾何構造と混合モデルの識別可能性構造に根本的に依存していることを確立した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。