Skip to main content
QUICK REVIEW

[论文解读] Derivatives and residual distribution of regularized M-estimators with application to adaptive tuning

Pierre Bellec, Yiwei Shen|arXiv (Cornell University)|Jul 11, 2021
Statistical Methods and Inference参考文献 17被引用 5
一句话总结

本文为高维线性模型中具有高斯设计和任意噪声的正则化M-估计器提出了一种新颖的自适应调参准则。通过推导M-估计器对响应变量和设计矩阵的雅可比矩阵的精确表达式,该研究刻画了残差分布,并提出了一种无需噪声分布或设计协方差信息即可近似预测误差的准则,从而实现了稳健且数据驱动的调参。

ABSTRACT

This paper studies M-estimators with gradient-Lipschitz loss function regularized with convex penalty in linear models with Gaussian design matrix and arbitrary noise distribution. A practical example is the robust M-estimator constructed with the Huber loss and the Elastic-Net penalty and the noise distribution has heavy-tails. Our main contributions are three-fold. (i) We provide general formulae for the derivatives of regularized M-estimators $\hatβ(y,X)$ where differentiation is taken with respect to both $y$ and $X$; this reveals a simple differentiability structure shared by all convex regularized M-estimators. (ii) Using these derivatives, we characterize the distribution of the residual $r_i = y_i-x_i^ op\hatβ$ in the intermediate high-dimensional regime where dimension and sample size are of the same order. (iii) Motivated by the distribution of the residuals, we propose a novel adaptive criterion to select tuning parameters of regularized M-estimators. The criterion approximates the out-of-sample error up to an additive constant independent of the estimator, so that minimizing the criterion provides a proxy for minimizing the out-of-sample error. The proposed adaptive criterion does not require the knowledge of the noise distribution or of the covariance of the design. Simulated data confirms the theoretical findings, regarding both the distribution of the residuals and the success of the criterion as a proxy of the out-of-sample error. Finally our results reveal new relationships between the derivatives of $\hatβ(y,X)$ and the effective degrees of freedom of the M-estimator, which are of independent interest.

研究动机与目标

  • 为在噪声分布未知且可能重尾的情况下,正则化M-估计器的调参选择提供理论基础明确的自适应准则。
  • 推导正则化M-估计器对响应向量y和设计矩阵X的导数的一般公式。
  • 刻画在样本量n ≈ p的中间高维情形下残差的分布。
  • 构建一种调参准则,该准则可近似出样本外预测误差(至多相差一个常数),且无需知晓噪声分布或设计协方差结构。

提出的方法

  • 在梯度-Lipschitz损失和凸惩罚下,推导M-估计器对y和X的雅可比矩阵,揭示了一种普遍的可微性结构。
  • 利用推导出的雅可比矩阵,通过矩阵V = diag(ψ′(r))(I - X(∂β̂/∂y))计算有效自由度(df)和残差协方差结构。
  • 提出调参准则 Crit(ρ,g) = ||r + (df̂ / tr[V])ψ(r)||²,该准则可近似出样本外误差(至多相差一个常数),其中r为残差向量,ψ为损失函数的导数。
  • 证明对于Huber损失与Elastic-Net等常见损失-惩罚对,比值df̂ / tr[V]具有闭式表达式,从而实现高效计算。
  • 通过在高斯和Rademacher设计下带有重尾噪声的模拟实验验证了该方法,结果确认了残差分布近似的准确性以及该准则的有效性。

实验结果

研究问题

  • RQ1在高维线性模型中,正则化M-估计器对响应向量和设计矩阵的导数的一般形式是什么?
  • RQ2在一般凸损失和惩罚下,当n ≈ p时,残差向量的分布如何?
  • RQ3能否构建一种无需噪声分布或设计协方差知识的数据驱动准则,以近似出样本外预测误差?
  • RQ4M-估计器的导数与其有效自由度之间存在何种关系?
  • RQ5当噪声为重尾或设计非高斯时,所提出的准则在调参选择方面表现如何?

主要发现

  • 正则化M-估计器对y的雅可比矩阵具有普遍结构,仅依赖于损失函数导数和惩罚函数的海森矩阵。
  • 残差向量r = y - Xβ̂的分布可在中间高维情形下被刻画,其中矩阵V和比值df̂ / tr[V]起着关键作用。
  • 所提出的准则 Crit(ρ,g) = ||r + (df̂ / tr[V])ψ(r)||² 可近似出样本外误差(至多相差一个与估计器无关的常数),使其成为模型选择的有效代理。
  • 该准则具有自适应性:无需知晓噪声分布或设计协方差Σ,且在重尾误差下表现良好。
  • 模拟结果证实,残差分布被理论模型良好近似,且该准则能成功选择使出样本外误差最小的调参。
  • 本研究揭示了M-估计器导数与有效自由度之间的新联系,其中比值df̂ / tr[V]为残差结构提供了关键连接。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。