Skip to main content
QUICK REVIEW

[论文解读] Finite-sample analysis of M-estimators using self-concordance

Dmitrii M. Ostrovskii, Francis Bach|arXiv (Cornell University)|Oct 16, 2018
Statistical Methods and Inference参考文献 57被引用 11
一句话总结

该论文通过自一致性质为M-估计量提供了有限样本界限,表明卡方型超额风险界限的临界样本量规模为 $ O(d \cdot d_{\text{eff}}) $,其中 $ d $ 为参数维度,$ d_{\text{eff}} $ 反映模型误设的影响。分析基于在总体最小化点处的局部次高斯性和曲率假设,当在Dikin椭球邻域内施加更强条件时,界限进一步收紧为 $ O(\max\{d_{\text{eff}}, d\log d\}) $,并在高斯设计下的逻辑回归中得到验证。

ABSTRACT

The classical asymptotic theory for parametric $M$-estimators guarantees that, in the limit of infinite sample size, the excess risk has a chi-square type distribution, even in the misspecified case. We demonstrate how self-concordance of the loss allows to characterize the critical sample size sufficient to guarantee a chi-square type in-probability bound for the excess risk. Specifically, we consider two classes of losses: (i) self-concordant losses in the classical sense of Nesterov and Nemirovski, i.e., whose third derivative is uniformly bounded with the $3/2$ power of the second derivative; (ii) pseudo self-concordant losses, for which the power is removed. These classes contain losses corresponding to several generalized linear models, including the logistic loss and pseudo-Huber losses. Our basic result under minimal assumptions bounds the critical sample size by $O(d \\cdot d_{\ ext{eff}}),$ where $d$ the parameter dimension and $d_{\ ext{eff}}$ the effective dimension that accounts for model misspecification. In contrast to the existing results, we only impose local assumptions that concern the population risk minimizer $\ heta_*$. Namely, we assume that the calibrated design, i.e., design scaled by the square root of the second derivative of the loss, is subgaussian at $\ heta_*$. Besides, for type-ii losses we require boundedness of a certain measure of curvature of the population risk at $\ heta_*$.Our improved result bounds the critical sample size from above as $O(\\max\\{d_{\ ext{eff}}, d \\log d\\})$ under slightly stronger assumptions. Namely, the local assumptions must hold in the neighborhood of $\ heta_*$ given by the Dikin ellipsoid of the population risk. Interestingly, we find that, for logistic regression with Gaussian design, there is no actual restriction of conditions: the subgaussian parameter and curvature measure remain near-constant over the Dikin ellipsoid. Finally, we extend some of these results to $\\ell_1$-penalized estimators in high dimensions.

研究动机与目标

  • 表征M-估计量在有限样本中实现卡方型超额风险界限的样本范围,超越渐近近似。
  • 利用损失函数的自一致性质,建立实现此类界限的样本量要求。
  • 通过仅在总体最小化点 $ \theta_* $ 处施加局部条件,放松全局光滑性假设。
  • 将结果扩展至高维设定下的$ \ell_1 $-惩罚估计量。
  • 在高斯设计下的逻辑回归中验证理论,表明在Dikin椭球范围内子高斯参数和曲率度量近乎恒定。

提出的方法

  • 引入两类损失函数:自一致损失(三阶导数受Hessian的3/2次方有界)与伪自一致损失(无此幂次约束)。
  • 通过将预测器按其二阶导数的平方根缩放,定义在 $ \theta_* $ 处的局部次高斯性,这是实现有限样本界限的关键假设。
  • 应用Dikin椭球邻域条件以加强界限,要求在 $ \theta_* $ 附近具备局部次高斯性和曲率控制。
  • 采用 $ \psi_2 $ 和 $ \psi_1 $-范数控制估计方程的尾部行为,从而实现矩和尾部界限分析。
  • 在自一致性质下,通过集中不等式推导超额风险界限,将样本量与有效维度 $ d_{\text{eff}} $ 联系起来。
  • 通过类似的局部假设和曲率控制,将结果扩展至 $ \ell_1 $-惩罚M-估计量。

实验结果

研究问题

  • RQ1M-估计量在有限样本中实现卡方型超额风险界限所需的临界样本量是多少?
  • RQ2损失函数的自一致性和伪自一致性如何影响有限样本收敛速率?
  • RQ3能否通过 $ \theta_* $ 处的局部假设(如校准预测器的次高斯性)替代全局光滑性条件?
  • RQ4Dikin椭球邻域条件是否能提供比全局假设更紧的样本量界限?
  • RQ5在 $ \ell_1 $-惩罚的高维设定下,这些界限的行为如何?

主要发现

  • 实现卡方型超额风险界限的临界样本量受 $ O(d \cdot d_{\text{eff}}) $ 限制,其中 $ d_{\text{eff}} $ 反映模型误设的影响。
  • 在更强的Dikin椭球邻域假设下,界限收紧为 $ O(\max\{d_{\text{eff}}, d\log d\}) $。
  • 在高斯设计下的逻辑回归中,子高斯参数和曲率度量在Dikin椭球范围内近乎恒定,验证了局部假设的合理性。
  • 该方法仅依赖于 $ \theta_* $ 处的局部正则性,避免了全局光滑性或曲率假设。
  • 分析可扩展至 $ \ell_1 $-惩罚M-估计量,在高维设定下提供有限样本保证。
  • 使用 $ \psi_1 $-范数可实现亚指数尾部控制,在模型正确设定时优于 $ \psi_2 $-范数界限。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。