[论文解读] Correction of overfitting bias in regression models
本文提出了一种基于自助法的新方法,用于校正当协变量数量 $p$ 与样本量 $n$ 相当时,回归模型中最大似然估计量的过拟合偏差。通过在正态性假设下从删一法理论推导出一组紧凑的非线性方程,该方法能够同时对回归参数和干扰参数进行偏差校正,实现了与各种 $\theta = p/n$ 制度下模拟结果高度一致的精确收缩因子。
Regression analysis based on many covariates is becoming increasingly common. However, when the number of covariates $p$ is of the same order as the number of observations $n$, maximum likelihood regression becomes unreliable due to overfitting. This typically leads to systematic estimation biases and increased estimator variances. It is crucial for inference and prediction to quantify these effects correctly. Several methods have been proposed in literature to overcome overfitting bias or adjust estimates. The vast majority of these focus on the regression parameters. But failure to estimate correctly also the nuisance parameters may lead to significant errors in confidence statements and outcome prediction. In this paper we present a jacknife method for deriving a compact set of non-linear equations which describe the statistical properties of the ML estimator in the regime where $p=O(n)$ and under the hypothesis of normally distributed covariates. These equations enable one to compute the overfitting bias of maximum likelihood (ML) estimators in parametric regression models as functions of $ζ= p/n$. We then use these equations to compute shrinkage factors in order to remove the overfitting bias of maximum likelihood (ML) estimators. This new derivation offers various benefits over the replica approach in terms of increased transparency and reduced assumptions. To illustrate the theory we performed simulation studies for multiple regression models. In all cases we find excellent agreement between theory and simulations.
研究动机与目标
- 解决当 $p = O(n)$ 时最大似然(ML)估计量的系统性偏差和方差膨胀问题,此情形下经典渐近理论失效。
- 将现有校正方法从回归系数扩展至包括残差方差等干扰参数,后者对置信区间和预测区间至关重要。
- 基于删一法理论,开发一种透明的、无需重复模拟的框架,用于推导最大似然估计量的渐近偏差和方差。
- 提供一种实用且计算高效的偏差校正方法,并通过模拟研究加以验证。
- 在医疗和生物统计应用中常见的高维设置下实现准确的统计推断,其中 $n$ 较小但 $p$ 较大。
提出的方法
- 从删一法分析中推导出一组非线性自洽方程,以描述在 $p = \theta n$ 制度下最大似然估计量的渐近行为。
- 利用协变量的旋转不变性和高斯假设,推导出估计量的随机表示,从而实现解析可处理性。
- 引入两个关键参数 $k_\star$ 和 $v_\star$,它们由方程计算得出,用于表征最大似然估计量的偏差和方差。
- 将这些参数应用于计算收缩因子,以同时校正 $\beta$ 和 $\sigma$ 估计量中的偏差与方差膨胀。
- 通过重叠集中度的几何解释,将 $k_\star$ 和 $v_\star$ 与可观测量联系起来。
- 通过模拟研究验证方法,比较不同 $\zeta = p/n$ 值下未校正与校正估计量的表现。
实验结果
研究问题
- RQ1当 $p$ 与 $n$ 同阶时,如何系统性地校正最大似然估计量的过拟合偏差?
- RQ2在高维回归中,回归系数与干扰参数(如残差方差)的联合渐近行为如何?
- RQ3能否开发一种无需重复模拟、透明的方法,以推导偏差与方差校正,而无需依赖统计物理假设?
- RQ4所推导的校正因子在有限样本设置下,能在多大程度上提高估计量精度并降低均方误差?
- RQ5估计量抽样分布的理论近似在不同 $p/n$ 比例下,与经验模拟结果的匹配程度如何?
主要发现
- 推导出的自洽方程在一系列 $\zeta = p/n$ 值下能准确预测最大似然估计量的偏差与方差,理论与模拟结果高度一致。
- 对 $\hat{\bm{\beta}}_n$ 和 $\hat{\bm{\sigma}}_n$ 的校正估计量基本无偏,其抽样分布中心位于真实参数值附近。
- 校正估计量 $\tilde{\bm{\beta}}_n$ 的近似渐近分布可良好地由 $\mathcal{N}(\bm{\beta}_0, v_\star^2 / k_\star^2 p)$ 描述,模拟直方图显示高度匹配。
- 该方法成功校正了在高维设置下 $\mathbf{e}_1^\prime\hat{\bm{\beta}}_n$ 方差被低估的问题,尤其当 $\bm{\beta}_0$ 稀疏时更为显著。
- 理论能准确捕捉 $\hat{\bm{\sigma}}_n$ 分布的众数,表明即使在过拟合情况下,对干扰参数的估计也可靠。
- 该方法对正态性假设的偏离具有鲁棒性,只要条件性中心极限定理适用于 $\mathbf{X}_i^\prime\hat{\bm{\beta}}_{(i)}$,表明其适用范围可超越高斯协变量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。