Skip to main content
QUICK REVIEW

[论文解读] Maximum Likelihood Estimation of Functionals of Discrete Distributions

Jiantao Jiao, Kartik Venkat|arXiv (Cornell University)|Jun 26, 2014
Statistical Methods and Inference被引用 7
一句话总结

本文对离散分布泛函的最大似然估计量(MLE)进行了非渐近分析,重点关注香农熵和幂和。通过结合浓度不等式与正线性算子的逼近理论,建立了最坏情况平方误差风险的紧致界,表明MLE在高维情形下严格次优,并揭示了其与最小最大最优估计量之间的样本复杂度差距。

ABSTRACT

We consider the problem of estimating functionals of discrete distributions, and focus on tight nonasymptotic analysis of the worst case squared error risk of widely used estimators. We apply concentration inequalities to analyze the random fluctuation of these estimators around their expectations, and the theory of approximation using positive linear operators to analyze the deviation of their expectations from the true functional, namely their \emph{bias}. We characterize the worst case squared error risk incurred by the Maximum Likelihood Estimator (MLE) in estimating the Shannon entropy $H(P) = \sum_{i = 1}^S -p_i \ln p_i$, and $F_α(P) = \sum_{i = 1}^S p_i^α,α>0$, up to multiplicative constants, for any alphabet size $S\leq \infty$ and sample size $n$ for which the risk may vanish. As a corollary, for Shannon entropy estimation, we show that it is necessary and sufficient to have $n \gg S$ observations for the MLE to be consistent. In addition, we establish that it is necessary and sufficient to consider $n \gg S^{1/α}$ samples for the MLE to consistently estimate $F_α(P), 0

研究动机与目标

  • 提供离散分布泛函MLE的最坏情况平方误差风险的非渐近分析。
  • 刻画MLE一致估计香农熵和幂和 $F_\alpha(P)$ 所需的样本复杂度。
  • 在高维情形($S \gg n$)下,比较MLE与最小最大率最优估计量的性能。
  • 评估狄利克雷先验平滑在熵估计中的有效性,表明其无法达到最小最大率。
  • 建立MLE的收敛速度在 $0 < \alpha < 1$ 和 $\alpha \in (1, 3/2)$ 时严格慢于最小最大最优率,除非 $\alpha \geq 3/2$。

提出的方法

  • 使用浓度不等式来控制MLE的随机波动与其期望之间的偏差。
  • 应用正线性算子的逼近理论分析MLE的偏差,即其期望与真实泛函之间的偏离。
  • 通过泰勒展开和二项分布计数的矩界分析 $F_\alpha(P) = \sum_i p_i^\alpha$ 的MLE方差。
  • 利用积分余项表示和尾概率估计推导 $\mathbb{E}[R_1(X;p)]$ 和 $\mathbb{E}[R_2(X;p)]$ 的界。
  • 采用不等式 $x^{2\alpha}e^{-cnx} \leq \left(\frac{2\alpha}{cen}\right)^{2\alpha}$ 来控制误差界中的指数衰减项。
  • 通过推导一致性所需的显式样本复杂度阈值,将MLE的风险与最小最大最优率进行比较。

实验结果

研究问题

  • RQ1MLE估计香农熵 $H(P)$ 的最坏情况平方误差风险是多少?其随样本量 $n$ 和字母表大小 $S$ 如何变化?
  • RQ2对于幂和 $F\alpha(P)$,MLE一致估计所需的必要且充分样本量 $n$ 是多少?其如何依赖于 $\alpha$?
  • RQ3当 $\alpha \in (0,1)$ 和 $\alpha \in (1, 3/2)$ 时,MLE的收敛速度与 $F_\alpha(P)$ 的最小最大最优率相比如何?
  • RQ4狄利克雷先验平滑能否改进熵估计中的MLE性能?其是否能达到最小最大率?
  • RQ5MLE与最小最大最优估计量在有效样本量方面的性能关系如何?

主要发现

  • 香农熵的MLE一致当且仅当 $n \gg S$,其最坏情况平方误差风险在绝对常数范围内是紧致的。
  • 对于 $0 < \alpha < 1$ 的 $F_\alpha(P)$,MLE需要 $n \gg S^{1/\alpha}$ 个样本才能一致,这严格多于最小最大最优的 $S^{1/\alpha}/\ln S$。
  • 当 $1 < \alpha < 3/2$ 时,MLE的最坏情况平方误差率是 $n^{-2(\alpha-1)}$,而最小最大率是 $(n\ln n)^{-2(\alpha-1)}$,显示出对数差距。
  • 当 $\alpha \geq 3/2$ 时,MLE的收敛速度为 $n^{-1}$,与字母表大小无关,达到最小最大最优率。
  • 使用 $n$ 个样本的最小最大率最优估计量的性能,与使用 $n\ln n$ 个样本的狄利克雷平滑估计量基本相当。
  • 无论参数如何调整,狄利克雷先验平滑(无论是插值法还是贝叶斯估计器)都无法在熵估计中达到最小最大率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。