Skip to main content
QUICK REVIEW

[论文解读] Nonparametric identification and maximum likelihood estimation for hidden Markov model

Grigory Alexandrovich, Hajo Holzmann|arXiv (Cornell University)|Apr 16, 2014
Bayesian Methods and Mixture Models参考文献 12被引用 6
一句话总结

本文在最小假设下建立了有限状态隐马尔可夫模型(HMM)的非参数识别性与非参数最大似然估计(NP-MLE)的一致性:即全秩、遍历的转移矩阵和各状态不同的状态依赖分布。证明了广义Kullback–Leibler散度唯一确定真实参数,并展示了当状态密度为任意参数族的混合时,即使不识别混合分布本身,NP-MLE仍具有一致性。

ABSTRACT

Nonparametric identification and maximum likelihood estimation for finite-state hidden Markov models are investigated. We obtain identification of the parameters as well as the order of the Markov chain if the transition probability matrices have full-rank and are ergodic, and if the state-dependent distributions are all distinct, but not necessarily linearly independent. Based on this identification result, we develop nonparametric maximum likelihood estimation theory. First, we show that the asymptotic contrast, the Kullback--Leibler divergence of the hidden Markov model, identifies the true parameter vector nonparametrically as well. Second, for classes of state-dependent densities which are arbitrary mixtures of a parametric family, we show consistency of the nonparametric maximum likelihood estimator. Here, identification of the mixing distributions need not be assumed. Numerical properties of the estimates as well as of nonparametric goodness of fit tests are investigated in a simulation study.

研究动机与目标

  • 在最小正则性条件下,建立HMM参数(包括状态数)的非参数识别性。
  • 证明广义Kullback–Leibler散度可作为HMM中NP-MLE的一致对比函数。
  • 证明当状态依赖密度为参数族的任意混合时,非参数最大似然估计具有一致性,即使不假设混合分布可识别。
  • 为HMM中的非参数推断提供理论基础,避免对分量密度的参数假设。

提出的方法

  • 利用时间反转法和Kruskal关于三阶数组的定理,建立给定隐状态的条件分布的识别性。
  • 应用Kingman的次可加遍历定理,证明渐近对比(广义KL散度)的存在性与正定性。
  • 依赖于给定隐状态的观测段条件分布函数的线性无关性,即使原始状态依赖分布本身线性相关。
  • 采用凸分析与有界Lipschitz度量,处理非参数混合类中混合测度的收敛性。
  • 构建单纯形值初始分布的联合分布,将渐近对比表示为普通KL散度的积分形式。
  • 使用高阶转移(阶数 $t_0 = K^2 - 2K + 2$)以确保转移矩阵与初始分布的正性,从而支持技术性论证。

实验结果

研究问题

  • RQ1在何种条件下,有限状态HMM的参数向量(包括状态数)可实现非参数识别?
  • RQ2在非参数设定下,广义Kullback–Leibler散度是否唯一确定真实HMM参数?
  • RQ3当状态依赖密度为参数族的任意混合时,即使混合分布不可识别,非参数最大似然估计是否仍具有一致性?
  • RQ4当状态依赖分布线性相关时(这在光滑或对数凹密度等非参数类中很常见),如何实现识别?

主要发现

  • 若转移矩阵为全秩且遍历,且各状态依赖分布互不相同,则有限状态HMM的参数(包括状态数)可实现非参数识别。
  • 广义Kullback–Leibler散度(定义为归一化对数似然的极限)是一种正定对比函数,能唯一确定真实HMM参数。
  • 对于状态依赖密度为参数族任意混合的分布类,即使不假设混合分布可识别,非参数最大似然估计仍具有一致性。
  • 通过使用时间反转HMM与Kruskal关于三阶数组的定理,克服了状态密度之间线性相关性的挑战。
  • 通过使用Lipschitz连续函数逼近密度点值,以及在混合测度上使用有界Lipschitz度量,建立了NP-MLE的一致性。
  • 证明了对于不同参数,渐近对比严格为正,从而确保真实模型是对比函数的唯一最小化者。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。