Skip to main content
QUICK REVIEW

[论文解读] High-Dimensional Quantile Regression: Convolution Smoothing and Concave Regularization

Kean Ming Tan, Lan Wang|arXiv (Cornell University)|Sep 12, 2021
Statistical Methods and Inference参考文献 10被引用 5
一句话总结

本文提出了一种卷积平滑的分位数回归方法,结合迭代重加权 $ε$-惩罚 $´1$-正则化,以克服标准分位数损失函数的非光滑性和缺乏强凸性问题。该方法在近乎必要且充分的最小信号强度条件下,实现了最优收敛速率和强oracle性质,从而在高维设置下实现一致的变量选择与高效估计。

ABSTRACT

$\ell_1$-penalized quantile regression is widely used for analyzing high-dimensional data with heterogeneity. It is now recognized that the $\ell_1$-penalty introduces non-negligible estimation bias, while a proper use of concave regularization may lead to estimators with refined convergence rates and oracle properties as the signal strengthens. Although folded concave penalized $M$-estimation with strongly convex loss functions have been well studied, the extant literature on quantile regression is relatively silent. The main difficulty is that the quantile loss is piecewise linear: it is non-smooth and has curvature concentrated at a single point. To overcome the lack of smoothness and strong convexity, we propose and study a convolution-type smoothed quantile regression with iteratively reweighted $\ell_1$-regularization. The resulting smoothed empirical loss is twice continuously differentiable and (provably) locally strongly convex with high probability. We show that the iteratively reweighted $\ell_1$-penalized smoothed quantile regression estimator, after a few iterations, achieves the optimal rate of convergence, and moreover, the oracle rate and the strong oracle property under an almost necessary and sufficient minimum signal strength condition. Extensive numerical studies corroborate our theoretical results.

研究动机与目标

  • 解决分位数回归损失函数的非光滑性与缺乏强凸性问题,以改善理论分析与估计效率。
  • 开发一种结合卷积平滑与迭代重加权 $´1$-正则化的算法,以实现强凸性并提升估计性能。
  • 为带有凹惩罚的高维分位数回归建立理论保证,特别是最优收敛速率与强oracle性质。
  • 推导出在估计器可实现oracle估计时,近乎必要且充分的最小信号强度条件。

提出的方法

  • 引入一种卷积型平滑方法,对分段线性的分位数损失函数进行平滑处理,得到一个二阶连续可微的样本损失函数。
  • 采用迭代重加权 $´1$-正则化方法近似凹惩罚,相比标准 $´1$-惩罚,可显著降低估计偏差。
  • 证明平滑后的损失函数以高概率具有局部强凸性,从而支持稳定优化。
  • 利用Bahadur表示法分析oracle估计量的渐近行为,并推导收敛速率。
  • 应用局部线性逼近框架,在适当的信号强度条件下,将平滑估计量与oracle解联系起来。
  • 利用浓度不等式与矩阵范数界控制估计误差,确保变量选择的一致性。

实验结果

研究问题

  • RQ1卷积平滑能否将非光滑的分位数损失转化为适合高维推断的强凸、可微函数?
  • RQ2在平滑损失下,采用迭代重加权 $´1$-正则化的分位数回归能否实现高维情况下的oracle性质?
  • RQ3为使估计器恢复真实非零变量集合并实现最优收敛速率,所需的最小信号强度条件是什么?
  • RQ4与标准 $´1$-惩罚分位数回归相比,平滑估计器在偏差与估计效率方面表现如何?
  • RQ5在何种条件下,平滑估计器可达到最优收敛速率与强oracle性质?

主要发现

  • 平滑后的样本损失函数是二阶连续可微的,且以高概率具有局部强凸性,支持稳定优化。
  • 经过几次迭代后,迭代重加权 $´1$-惩罚的平滑分位数回归估计量可实现最优收敛速率。
  • 在近乎必要且充分的最小信号强度条件下,估计量可实现强oracle性质。
  • 参数估计量的收敛速率为 $O_p(\sqrt{s \log p / n})$,与经典稀疏估计中的最优速率一致。
  • 当最小非零系数满足 $\|\bm{\beta}^*_{\mathcal{S}}\|_{\min} \gtrsim \sqrt{s \log p / n}$ 时,估计量以高概率一致选择出真实活跃集。
  • 数值实验验证了理论结果,表明该方法在变量选择与估计精度方面优于标准 $´1$-惩罚分位数回归。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。