Skip to main content
QUICK REVIEW

[论文解读] A comparative study of new cross-validated bandwidth selectors for kernel density estimation

Enno Mammen, Miranda, Maria Dolores Martinez|arXiv (Cornell University)|Sep 20, 2012
Sparse and Compressive Sensing Techniques参考文献 11被引用 4
一句话总结

本文评估了核密度估计中新型交叉验证带宽选择器,重点关注间接交叉验证与多项式及高斯核的do-validation。研究发现,尽管高阶多项式间接核在间接交叉验证中理论上表现更优,但在do-validation中这一优势并不成立;实际中,原始do-validation仍更优,因其偏差更低,在有限样本中甚至优于组合估计器与插值法。

ABSTRACT

Recent contributions to kernel smoothing show that the performance of cross-validated bandwidth selectors improve significantly from indirectness. Indirect crossvalidation first estimates the classical cross-validated bandwidth from a more rough and difficult smoothing problem than the original one and then rescales this indirect bandwidth to become a bandwidth of the original problem. The motivation for this approach comes from the observation that classical crossvalidation tends to work better when the smoothing problem is difficult. In this paper we find that the performance of indirect crossvalidation improves theoretically and practically when the polynomial order of the indirect kernel increases, with the Gaussian kernel as limiting kernel when the polynomial order goes to infinity. These theoretical and practical results support the often proposed choice of the Gaussian kernel as indirect kernel. However, for do-validation our study shows a discrepancy between asymptotic theory and practical performance. As for indirect crossvalidation, in asymptotic theory the performance of indirect do-validation improves with increasing polynomial order of the used indirect kernel. But this theoretical improvements do not carry over to practice and the original do-validation still seems to be our preferred bandwidth selector. We also consider plug-in estimation and combinations of plug-in bandwidths and crossvalidated bandwidths. These latter bandwidths do not outperform the original do-validation estimator either.

研究动机与目标

  • 评估新型交叉验证带宽选择器(特别是间接交叉验证与do-validation)在核密度估计中的性能。
  • 探究提高间接核多项式阶数是否能在理论上与实践中改善带宽选择。
  • 将do-validation、插值法与混合估计器的性能与经典交叉验证及最优带宽进行比较。
  • 确定带宽选择器的理论改进是否能在有限样本设置中转化为实际性能提升。

提出的方法

  • 研究采用间接交叉验证,即先为一个更困难的平滑问题选择带宽(使用不同核),再将其重新缩放用于原始问题。
  • 评估多项式阶数的间接核与极限高斯核,分析其理论性能与有限样本表现。
  • 将do-validation作为具有改进渐近性质的间接交叉验证方法应用,并比较其在不同间接核阶数下的表现。
  • 通过组合8个交叉验证带宽(来自4种间接核与4种do-validation变体)和5个插值带宽,构建中位数估计器,并将中位数作为最终估计。
  • 在样本量n = 100、200、500与1000下,针对六种密度设计进行模拟,使用均方积分误差(MISE)与第90百分位数ISE作为性能指标。
  • 所有方法的性能均以理论ISE最优带宽为基准,并通过经验分位数评估,以反映实际应用中的可靠性。

实验结果

研究问题

  • RQ1提高间接核的多项式阶数是否能改善间接交叉验证的实际性能?
  • RQ2为何do-validation在使用高阶间接核时的理论改进无法转化为有限样本中的更好表现?
  • RQ3结合多个带宽选择器的中位数估计器,在鲁棒性与准确性方面与单一方法相比如何?
  • RQ4尽管具有理论优势,插值法或混合估计器是否能在实际设置中超越原始do-validation?
  • RQ5在应用带宽选择中,使用ISE的第90百分位数是否比均值ISE更具参考价值?

主要发现

  • 间接交叉验证的性能在理论上与实践中均随间接核多项式阶数的提高而改善,高斯核为最优极限。
  • 对于do-validation,高阶间接核带来的理论改进无法在实际中体现;原始do-validation因偏差更低而保持更优表现。
  • 中位数估计器(结合8个交叉验证带宽与5个插值带宽)表现接近do-validation,但在简洁性与计算效率方面表现较差。
  • 插值法虽理论性能更优,但偏差较高,在所有有限样本性能指标中均被do-validation超越。
  • 当使用ISE的第90百分位数作为性能指标时,do-validation与中位数估计器的优势更加显著,表明其在实际应用中更具可靠性。
  • 本研究证实,实际实现中的挑战(如偏差累积)往往超过带宽选择器设计中的理论收益。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。