Skip to main content
QUICK REVIEW

[论文解读] Non-asymptotic Theory for the Plug-in Rule in Functional Estimation

Jiantao Jiao, Kartik Venkat|arXiv (Cornell University)|Jun 26, 2014
Statistical Methods and Inference参考文献 78被引用 6
一句话总结

该论文为函数估计中的插补规则发展了一套非渐近理论,利用正线性算子的逼近理论分析偏差和方差的集中不等式。研究结果表明,最大似然估计(MLE)在估计微分熵和R{\alpha}-泛函时严格次优,其极小极大样本复杂度比MLE性能高出对数因子。

ABSTRACT

The plug-in rule is widely used in estimating functionals of finite dimensional parameters, and involves plugging in an asymptotically efficient estimator for the parameter to obtain an asymptotically efficient estimator for the functional. We propose a general non-asymptotic theory for analyzing the performance of the plug-in rule, and demonstrate its utility by applying it to estimation of functionals of discrete distributions via the maximum likelihood estimator (MLE). We show that existing theory is insufficient for analyzing the bias of the plug-in rule, and propose to apply the theory of approximation using positive linear operators to study this bias. The variance is controlled using the well-known tools from the literature on concentration inequalities. Our techniques completely characterize the maximum $L_2$ risk incurred by the MLE in estimating the Shannon entropy $H(P) = \sum_{i = 1}^S -p_i \ln p_i$, and $F_\alpha(P) = \sum_{i = 1}^S p_i^\alpha$ up to a constant. As corollaries, for Shannon entropy estimation, we show that it is necessary and sufficient to have $n = \omega(S)$ observations for the MLE to be consistent, where $S$ represents the alphabet size. In addition, we obtain that it is necessary and sufficient to consider $n = \omega(S^{1/\alpha})$ samples for the MLE to consistently estimate $F_\alpha(P), 0<\alpha<1$. The minimax sample complexity for both problems are $\omega(S/\ln S)$ and $\omega(S^{1/\alpha}/\ln S)$, which implies that the MLE is strictly sub-optimal. When $1<\alpha<3/2$, we show that the maximum $L_2$ rate of convergence for the MLE is $n^{-2(\alpha-1)}$ for infinite alphabet size, while the minimax $L_2$ rate is $(n\ln n)^{-2(\alpha-1)}$. When $\alpha\geq 3/2$, the MLE achieves the minimax optimal $L_2$ convergence rate $n^{-1}$ regardless of the alphabet size.

研究动机与目标

  • 开发一个用于分析函数估计中插补规则的非渐近框架,特别是针对离散分布。
  • 解决现有渐近理论在刻画插补估计量偏差方面的不足。
  • 为估计香农熵和R{\alpha}-泛函等泛函的$L_2$风险提供精确的有限样本界。
  • 确定极小极大样本复杂度,并与MLE的性能进行比较。
  • 建立MLE在估计熵和$F_\alpha(P)$时一致性的样本量$n$的必要与充分条件。

提出的方法

  • 应用正线性算子的逼近理论来建模和控制插补估计量的偏差。
  • 使用集中不等式来限制有限样本中插补估计量的方差。
  • 刻画MLE在估计香农熵$H(P)$和$F_\alpha(P) = \sum p_i^\alpha$时的最大$L_2$风险。
  • 推导$L_2$风险的非渐近上下界,以建立极小极大最优性。
  • 分析$L_2$风险对字母表大小$S$和参数$\alpha$的依赖关系。
  • 在不同$\alpha$的参数范围内,比较MLE的收敛速率与极小极大最优速率。

实验结果

研究问题

  • RQ1当使用MLE作为参数估计器时,插补规则的非渐近偏差行为如何?
  • RQ2MLE在估计香农熵时的$L_2$风险如何随字母表大小$S$和样本量$n$变化?
  • RQ3估计$F_\alpha(P)$所需的极小极大样本复杂度是多少?与MLE性能相比如何?
  • RQ4在哪些$\alpha$值下,MLE在$L_2$风险下是极小极大最优的?在哪些情况下是严格次优的?
  • RQ5当$\alpha \geq 3/2$时,MLE对$F_\alpha(P)$的收敛速率是多少?与极小极大速率相比如何?

主要发现

  • MLE在估计香农熵时严格次优,其极小极大样本复杂度为$\omega(S / \ln S)$,超过MLE所需的$\omega(S)$。
  • 对于$0 < \alpha < 1$的$F_\alpha(P)$,极小极大样本复杂度为$\omega(S^{1/\alpha} / \ln S)$,而MLE仅需$\omega(S^{1\alpha})$,显示出对数差距。
  • 当$1 < \alpha < 3/2$时,MLE的$L_2$风险速率为$n^{-2(\alpha - 1)}$,严格慢于极小极大速率$(n \ln n)^{-2(\alpha - 1)}$。
  • 当$\alpha \geq 3/2$时,MLE在$L_2$风险下达到极小极大最优速率$n^{-1}$,与字母表大小无关。
  • 通过所提出的非渐近框架,MLE在估计香农熵和$F_\alpha(P)$时的最大$L_2$风险被完全刻画,仅相差一个常数因子。
  • MLE在估计香农熵时一致性的必要与充分条件是$n = \omega(S)$;对于$0 < \alpha < 1$的$F_\alpha(P)$,必要与充分条件是$n = \omega(S^{1/\alpha})$。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。