Skip to main content
QUICK REVIEW

[论文解读] Convergence rates of least squares regression estimators with heavy-tailed errors

Qiyang Han, Jon A. Wellner|arXiv (Cornell University)|Jun 7, 2017
Statistical Methods and Inference参考文献 58被引用 5
一句话总结

本文在仅具有 $p$-阶矩($p \geq 1$)的重尾误差下,建立了非参数回归中最小二乘估计量(LSE)的收敛速率。在熵条件指数 $\alpha \in (0,2)$ 下,LSE 的 $L_2$ 预测风险以速率 $\mathcal{O}_\mathbf{P}(n^{-1/(2+\alpha)} \vee n^{-1/2 + 1/(2p)})$ 收敛,当 $p \geq 1 + 2/\alpha$ 时与高斯分布下的速率一致,但当 $p < 1 + 2/\alpha$ 时慢于鲁棒估计量。关键技术进展是为乘子经验过程提出了一项新型乘子不等式,从而在重尾条件下实现了精确的界控制。

ABSTRACT

We study the performance of the Least Squares Estimator (LSE) in a general nonparametric regression model, when the errors are independent of the covariates but may only have a $p$-th moment ($p\geq 1$). In such a heavy-tailed regression setting, we show that if the model satisfies a standard `entropy condition' with exponent $α\in (0,2)$, then the $L_2$ loss of the LSE converges at a rate \begin{align*} \mathcal{O}_{\mathbf{P}}\big(n^{-\frac{1}{2+α}} \vee n^{-\frac{1}{2}+\frac{1}{2p}}\big). \end{align*} Such a rate cannot be improved under the entropy condition alone. This rate quantifies both some positive and negative aspects of the LSE in a heavy-tailed regression setting. On the positive side, as long as the errors have $p\geq 1+2/α$ moments, the $L_2$ loss of the LSE converges at the same rate as if the errors are Gaussian. On the negative side, if $p&lt;1+2/α$, there are (many) hard models at any entropy level $α$ for which the $L_2$ loss of the LSE converges at a strictly slower rate than other robust estimators. The validity of the above rate relies crucially on the independence of the covariates and the errors. In fact, the $L_2$ loss of the LSE can converge arbitrarily slowly when the independence fails. The key technical ingredient is a new multiplier inequality that gives sharp bounds for the `multiplier empirical process' associated with the LSE. We further give an application to the sparse linear regression model with heavy-tailed covariates and errors to demonstrate the scope of this new inequality.

研究动机与目标

  • 研究当误差仅具有 $p$-阶矩($p \geq 1$)而非次高斯或次指数尾部时,非参数回归中最小二乘估计量(LSE)的收敛行为。
  • 确定在何种条件下,LSE 能够达到与高斯误差下相同的收敛速率,以及在何种条件下其性能被鲁棒估计量超越。
  • 发展一项新的技术工具——针对重尾误差下乘子经验过程的精确乘子不等式,该工具是推导收敛速率的核心。
  • 建立LSE收敛速率取决于函数类的熵指数 $\alpha$ 与误差的矩 $p$ 之间相互作用的结论,临界阈值为 $p = 1 + 2/\alpha$。

提出的方法

  • 推导出一项新型乘子不等式,为重尾误差下的乘子经验过程提供精确界,从而控制LSE风险分解中的随机项。
  • 结合熵积分法与对称化技巧,界定了由函数类 $\mathcal{F}$ 索引的经验过程的期望上确界。
  • 采用覆盖数条件 $\log \mathcal{N}(\varepsilon, \mathcal{F}, L_2(Q)) \lesssim \varepsilon^{-\alpha}$($\alpha \in (0,2)$)的链式论证,量化了函数类的复杂度。
  • 通过平衡熵条件带来的偏差与误差的 $p$-阶矩带来的方差,建立了 $L_2$ 预测风险 $\|\hat{f}_n - f_0\|_{L_2(P)}$ 的精确速率。
  • 将主不等式应用于具有重尾协变量和误差的稀疏线性回归模型,展示了该方法的广泛适用性。
  • 采用对称化与去对称化技术,将误差 $\xi_i$ 的经验过程期望与涉及 $\varepsilon_i$ 的 Rademacher 混合型项关联起来,从而实现矩界控制。

实验结果

研究问题

  • RQ1在误差仅具有 $p$-阶矩($p \geq 1$)的非参数回归中,是什么决定了LSE的收敛速率?
  • RQ2在何种条件下,LSE 能够达到与次高斯误差下相同的速率,而在何种条件下无法实现?
  • RQ3函数类的熵指数 $\alpha$ 与误差的矩 $p$ 之间的相互作用如何影响LSE的收敛速率?
  • RQ4能否发展出一项新的乘子不等式以处理重尾误差下的乘子经验过程,其对LSE风险界具有何种影响?
  • RQ5所推导的收敛速率 $\mathcal{O}_\mathbf{P}(n^{-1/(2+\alpha)} \vee n^{-1/2 + 1/(2p)})$ 是否为精确速率,且仅在熵条件下能否进一步改进?

主要发现

  • 在熵条件指数 $\alpha \in (0,2)$ 与 $p$-阶矩误差下,LSE 的 $L_2$ 预测风险以速率 $\mathcal{O}_\mathbf{P}(n^{-1/(2+\alpha)} \vee n^{-1/2 + 1/(2p)})$ 收敛。
  • 当 $p \geq 1 + 2/\alpha$ 时,LSE 达到与高斯误差下相同的速率 $n^{-1/(2+\alpha)}$,表明对重尾具有鲁棒性。
  • 当 $p < 1 + 2/\alpha$ 时,LSE 的收敛速度严格慢于鲁棒估计量,且速率 $n^{-1/2 + 1/(2p)}$ 在任意熵水平 $\alpha$ 下对某些模型均为不可避免。
  • 所推导的速率是精确的,仅在熵条件下无法进一步改进,通过构造达到该速率的困难模型予以证明。
  • 关键技术贡献——新型乘子不等式——为重尾误差下的乘子经验过程提供了精确界,从而支持了速率分析。
  • 协变量与误差的独立性至关重要:若无此假设,LSE 的 $L_2$ 损失可能收敛至任意缓慢,导致所推导的速率不成立。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。