Skip to main content
QUICK REVIEW

[论文解读] Nonregular and Minimax Estimation of Individualized Thresholds in High Dimension with Binary Responses

Huijie Feng, Yang Ning|arXiv (Cornell University)|May 26, 2019
Statistical Methods and Inference参考文献 37被引用 4
一句话总结

本文提出了一种正则化平滑损失方法,用于高维二值响应模型中个体化线性阈值的估计,解决了非正则性和计算不可行性问题。该方法建立了非标准的误差率 $$(s\log d/n)^{\beta/(2\beta+1)}$$,其最优性在对数因子范围内达到极小极大最优,并通过自适应 Lepski 方法与路径追踪算法确保了几何收敛速度。

ABSTRACT

Given a large number of covariates $Z$, we consider the estimation of a high-dimensional parameter $θ$ in an individualized linear threshold $θ^T Z$ for a continuous variable $X$, which minimizes the disagreement between $ ext{sign}(X-θ^TZ)$ and a binary response $Y$. While the problem can be formulated into the M-estimation framework, minimizing the corresponding empirical risk function is computationally intractable due to discontinuity of the sign function. Moreover, estimating $θ$ even in the fixed-dimensional setting is known as a nonregular problem leading to nonstandard asymptotic theory. To tackle the computational and theoretical challenges in the estimation of the high-dimensional parameter $θ$, we propose an empirical risk minimization approach based on a regularized smoothed loss function. The statistical and computational trade-off of the algorithm is investigated. Statistically, we show that the finite sample error bound for estimating $θ$ in $\ell_2$ norm is $(s\log d/n)^{β/(2β+1)}$, where $d$ is the dimension of $θ$, $s$ is the sparsity level, $n$ is the sample size and $β$ is the smoothness of the conditional density of $X$ given the response $Y$ and the covariates $Z$. The convergence rate is nonstandard and slower than that in the classical Lasso problems. Furthermore, we prove that the resulting estimator is minimax rate optimal up to a logarithmic factor. The Lepski's method is developed to achieve the adaption to the unknown sparsity $s$ and smoothness $β$. Computationally, an efficient path-following algorithm is proposed to compute the solution path. We show that this algorithm achieves geometric rate of convergence for computing the whole path. Finally, we evaluate the finite sample performance of the proposed estimator in simulation studies and a real data analysis.

研究动机与目标

  • 解决高维二值响应个体化阈值估计中的计算不可行性与非正则渐近行为问题。
  • 为线性阈值模型 $\bm{\theta}^T\bm{Z}$ 中高维参数 $\bm{\theta}$ 的估计,开发一种统计最优且计算高效的算法。
  • 在未知稀疏性 $s$ 与光滑性 $\beta$ 的条件下,建立估计量的有限样本误差界与极小极大最优性。
  • 通过 Lepski 方法实现对未知 $s$ 与 $\beta$ 的自适应,并确保计算过程的几何收敛。

提出的方法

  • 将阈值估计建模为 M-估计问题,通过最小化带不连续符号函数的经验风险。
  • 引入正则化平滑损失函数以替代不可微的符号函数,从而实现可计算的优化。
  • 证明当平滑带宽 $\delta \to 0$ 时具有 Fisher 一致性,确保方法收敛至真实风险最小化器。
  • 在给定 $Y$ 与 $\bm{Z}$ 条件下,$X$ 的条件密度具有光滑性 $\beta$ 的假设下,推导出 $\ell_2$ 误差界为 $(s\log d/n)^{\beta/(2\beta+1)}$。
  • 应用 Lepski 方法自适应选择调参,无需事先知晓 $s$ 与 $\beta$。
  • 开发一种路径追踪算法,实现全解路径的几何收敛速率,确保高效计算。

实验结果

研究问题

  • RQ1能否在非正则条件下,为高维二值响应阈值估计开发一种计算可行且统计最优的方法?
  • RQ2在未知光滑性 $\beta$ 与稀疏性 $s$ 的条件下,$\bm{\theta}$ 在 $\ell_2$ 范数下的最优收敛速率为何?
  • RQ3如何在高维阈值估计中实现对未知 $s$ 与 $\beta$ 的自适应?
  • RQ4能否通过正则化平滑损失方法有效求解非凸、非光滑的经验风险最小化问题?
  • RQ5所提方法在高维、非正则设定下是否能达到极小极大最优性(对数因子范围内)?

主要发现

  • 所提估计量实现了 $(s\log d/n)^{\beta/(2\beta+1)}$ 的 $\ell_2$ 误差界,该结果为非标准形式,且慢于经典 Lasso 的收敛速率。
  • 该收敛速率在对数因子范围内达到极小极大最优,确立了该方法的理论最优性。
  • Lepski 方法实现了对未知稀疏性 $s$ 与光滑性 $\beta$ 的自适应,无需预先知晓这些参数。
  • 路径追踪算法实现了几何收敛速率,确保全解路径的高效计算。
  • 模拟研究与 ChAMP 试验的真实数据分析验证了该方法在有限样本下的性能与鲁棒性,适用于不同变量选择模式。
  • 与 SVM 和逻辑回归等替代方法相比,该方法在变量选择方面表现更优,临床相关变量如 KSymp_3mo 与 SF36Soc_6mo 均稳定地被选中且系数为负。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。