[论文解读] Robust high dimensional learning for Lipschitz and convex losses
本文通过正则化经验风险最小化(RERM)和一种新颖的极小极大 MOM 估计器,发展了针对利普希茨连续和凸损失函数的鲁棒高维学习方法。在放宽的矩条件假设下,建立了最优次高斯偏差界,使方法在重尾分布和含异常值的数据设置中仍能保持鲁棒性能,适用于 LASSO、SLOPE、组 LASSO 和总变差正则化等场景。
We establish risk bounds for Regularized Empirical Risk Minimizers (RERM) when the loss is Lipschitz and convex and the regularization function is a norm. In a first part, we obtain these results in the i.i.d. setup under subgaussian assumptions on the design. In a second part, a more general framework where the design might have heavier tails and data may be corrupted by outliers both in the design and the response variables is considered. In this situation, RERM performs poorly in general. We analyse an alternative procedure based on median-of-means principles and called minmax MOM. We show optimal subgaussian deviation rates for these estimators in the relaxed setting. The main results are meta-theorems allowing a wide-range of applications to various problems in learning theory. To show a non-exhaustive sample of these potential applications, it is applied to classification problems with logistic loss functions regularized by LASSO and SLOPE, to regression problems with Huber loss regularized by Group LASSO and Total Variation. Another advantage of the minmax MOM formulation is that it suggests a systematic way to slightly modify descent based algorithms used in high-dimensional statistics to make them robust to outliers. We illustrate this principle in a Simulations section where a minmax MOM version of classical proximal descent algorithms are turned into robust to outliers algorithms.
研究动机与目标
- 将正则化经验风险最小化器(RERM)的鲁棒风险界扩展至具有利普希茨和凸损失的高维设置。
- 通过提出一种极小极大 MOM 估计器,解决在重尾设计或数据污染下 RERM 的不稳定性问题。
- 推导元定理,使其可广泛应用于各类结构化估计问题,包括 LASSO、SLOPE、组 LASSO 和总变差正则化。
- 提出一种系统化框架,通过极小极大 MOM 原理将标准下降算法转化为鲁棒版本。
- 在放宽的矩条件假设下,建立最优次高斯偏差速率,以矩条件替代对设计矩阵的次高斯假设。
提出的方法
- 基于分位数中位数原理,提出一种极小极大 MOM 估计器,以提升对异常值和重尾数据的鲁棒性。
- 引入局部复杂度参数和局部伯恩斯坦条件,以放宽先前工作中使用的全局假设。
- 通过在分块上最小化最大经验风险,将极小极大 MOM 框架应用于正则化估计器,确保稳定性。
- 利用 McDiarmid 不等式、Giné-Zinn 对称化方法和压缩论证,推导出 MOM 估计器偏差的高概率界。
- 通过稀疏性方程和分块分析推导风险界,确保在矩条件假设下估计器可达到最优收敛速率。
- 提出一种通用方法,通过将经验风险替换为极小极大 MOM 风险,将标准的近端下降算法转化为鲁棒版本。
实验结果
研究问题
- RQ1RERM 是否能在设计矩阵的矩假设下(而非次高斯假设)实现最优次高斯偏差界?
- RQ2如何构建极小极大 MOM 估计器,以确保对设计变量和响应变量中异常值的鲁棒性?
- RQ3局部伯恩斯坦条件在高维设置下推导鲁棒估计器快速收敛速率中起什么作用?
- RQ4极小极大 MOM 框架在 LASSO、SLOPE、组 LASSO 和总变差等结构化正则化问题中的适用范围有多大?
- RQ5标准下降算法是否可系统性地通过极小极大 MOM 原理改造为鲁棒版本?
主要发现
- 在矩假设下,即使设计矩阵具有重尾分布或数据受异常值污染,极小极大 MOM 估计器仍能实现最优次高斯偏差界。
- 所提出的元定理允许风险界依赖于局部复杂度参数,优于全局复杂度度量。
- 对于极小极大 MOM 估计器,事件 $ar{ig{ heta}}_K$ 以至少 $1 - 2 ext{exp}(-cK)$ 的概率成立,确保高概率稳定性。
- 估计器以高概率满足 $P ilde{ ho}^2 ilde{ ho}^2( ho^*, 2 ho^*)$,表明其具有快速收敛速率。
- 该框架可在最小假设下推导出 LASSO、SLOPE、组 LASSO 和总变差正则化问题的精确风险界。
- 极小极大 MOM 原理提供了一种系统化方法,可将标准近端下降算法转化为鲁棒版本,仿真结果已验证其有效性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。