Skip to main content
QUICK REVIEW

[论文解读] Robust and efficient estimation of high dimensional scatter and location

Ricardo A. Maronna, Vı́ctor J. Yohai|arXiv (Cornell University)|Apr 13, 2015
Advanced Statistical Methods and Models参考文献 21被引用 5
一句话总结

本文提出了高维多变量位置与散点矩阵的稳健且高效的估计量,重点研究了具有非单调权重函数(Rocke型)的S-估计量和MM-估计量。研究结果表明,经Peña和Prieto提出的半确定性初始化程序启动的Rocke估计量,在 $ p \geq 15 $ 时可实现高效率与高稳健性;而MM-估计量在 $ p < 15 $ 时表现最优,其效率与抗 outlier 能力均优于传统的MCD和MVE估计量。

ABSTRACT

We deal with the equivariant estimation of scatter and location for p-dimensional data, giving emphasis to scatter. It it important that the estimators possess both a high efficiency for normal data and a high resistance to outliers, that is, a low bias under contamination. The most frequently employed estimators are not quite satisfactory in this respect. The Minimum Volume Ellipsoid (MVE) and Minimum Covariance Determinant (MCD) estimators are known to have a very low efficiency. S-Estimators (Davies 1987) with a monotonic weight function like the bisquare behave satisfactorily for "small" p, say p not larger than 10. Rocke (1996) showed that their efficiency tends to one with increasing p. Unfortunately, this advantage is paid with a serious loss of robustness for large p. We consider three families of estimators with controllable efficiencies: non-monotonic S-estimators (Rocke 1996), MM-estimators (Tatsuoka and Tyler 2000) and tau-estimators (Lopuhaa 1991), whose performance for large p has not been explored to date. Two types of starting estimators are employed: the MVE computed through subsampling, and a semi-deterministic procedure proposed by Peña and Prieto (2007) for outlier detection. A simulation study shows that the Rocke and MM estimators starting from the Peña-Prieto estimator and with an adequate tuning, can simultaneously attain high efficiency and high robustness.

研究动机与目标

  • 为解决现有等变估计量在高维情形下的局限性,即传统方法(如MCD和MVE)效率低下,或具有单调权重的S-估计量在维度增加时稳健性下降的问题。
  • 在高维污染情景下,评估并比较四类估计量——非单调S-估计量(Rocke型)、MM-估计量、$\tau$-估计量和Stahel-Donoho估计量。
  • 评估起始估计量对计算效率与统计性能的影响,特别是对比子抽样与半确定性Peña-Prieto方法的差异。
  • 识别能同时最大化正态数据下效率与污染下稳健性的最优调参常数与权重函数。
  • 基于维度 $ p $ 建立经验性指导原则,推荐 $ p \geq 15 $ 时使用Rocke估计量,$ p < 15 $ 时使用MM-估计量。

提出的方法

  • 在S-估计中采用非单调权重函数,特别是Rocke(1996)的方法,以在保持稳健性的同时提升高维下的效率。
  • 采用MM-估计量,结合高 breakdown 初始估计与高效率重加权,以马氏距离的稳健尺度为基础。
  • 应用Peña和Prieto(2007)提出的半确定性起始程序进行异常值检测,相较于子抽样,可显著提升计算速度与估计精度。
  • 使用Kullback-Leibler散度作为污染下期望偏差的度量,以补充破缺点(breakdown point)的稳健性评估。
  • 对Stahel-Donoho估计量采用有限方向集,使用硬拒绝权重函数 $ W(t) = \mathbf{1}(t \leq \beta) $,并聚焦于基于峰度的方向选择方法(KSD变体)。
  • 通过全面的模拟研究,比较不同维度 $ p $、污染水平与样本大小下的各类估计量,采用有限样本效率与稳健性指标作为评估标准。

实验结果

研究问题

  • RQ1非单调S-估计量(Rocke型)是否能在高维情形($ p \geq 15 $)下同时实现高效率与高稳健性?
  • RQ2起始估计量的选择(子抽样 vs. Peña-Prieto方法)如何影响稳健散点与位置估计量的性能?
  • RQ3MM-估计量在低维数据($ p < 15 $)下是否在效率与稳健性方面优于其他等变估计量?
  • RQ4在椭球分布下,KSD估计量在点质量污染下的理论稳健性(以Kullback-Leibler散度与破缺点衡量)如何?
  • RQ5能否选择调参常数,使估计量在不同维度下同时实现正态数据下的高效率与污染下的低偏差?

主要发现

  • 经Peña-Prieto程序初始化并合理调参的Rocke估计量,在 $ p \geq 15 $ 时可同时实现高效率与高稳健性,优于单调S-估计量与MCD估计量。
  • 对于 $ p < 15 $,推荐使用MM-估计量,因其在效率与稳健性方面优于其他方法,包括MCD与MVE。
  • 半确定性Peña-Prieto起始程序相比标准子抽样,能显著降低计算时间并提升估计精度,尤其适用于MCD及其他迭代估计量。
  • 采用单调权重函数(如bisquare)的S-估计量在小 $ p $ 时效率较低;尽管其效率随 $ p $ 增加而上升,但在高维下会丧失稳健性。
  • 理论分析表明,采用基于峰度的方向选择与硬拒绝权重的KSD估计量,在点质量污染下破缺点可达0.5,前提是其单变量估计量 $ \mu $ 与 $ \sigma $ 本身也具有0.5的破缺点。
  • 模拟研究证实,采用Peña-Prieto起始的Rocke估计量在 $ p \geq 15 $ 时,即使样本量适中,也能保持低偏差与高效率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。