Skip to main content
QUICK REVIEW

[论文解读] Excess risk bounds in robust empirical risk minimization

Stanislav Minsker, Timothée Mathieu|arXiv (Cornell University)|Oct 16, 2019
Statistical Methods and Inference参考文献 43被引用 8
一句话总结

该论文提出了一种鲁棒的经验风险最小化框架,用鲁棒估计量(特别是分位数均值和Catoni型估计量)替代标准样本均值,以处理重尾分布和对抗性污染。该框架在仅需极低矩假设(如有限二阶至四阶矩)下,建立了高概率的超额风险界,且收敛速率较快,明确依赖于样本大小和污染水平。

ABSTRACT

This paper investigates robust versions of the general empirical risk minimization algorithm, one of the core techniques underlying modern statistical methods. Success of the empirical risk minimization is based on the fact that for a "well-behaved" stochastic process $\left\{ f(X), \ f\in \mathcal F ight\}$ indexed by a class of functions $f\in \mathcal F$, averages $\frac{1}{N}\sum_{j=1}^N f(X_j)$ evaluated over a sample $X_1,\ldots,X_N$ of i.i.d. copies of $X$ provide good approximation to the expectations $\mathbb E f(X)$ uniformly over large classes $f\in \mathcal F$. However, this might no longer be true if the marginal distributions of the process are heavy-tailed or if the sample contains outliers. We propose a version of empirical risk minimization based on the idea of replacing sample averages by robust proxies of the expectation, and obtain high-confidence bounds for the excess risk of resulting estimators. In particular, we show that the excess risk of robust estimators can converge to $0$ at fast rates with respect to the sample size. We discuss implications of the main results to the linear and logistic regression problems, and evaluate the numerical performance of proposed methods on simulated and real data.

研究动机与目标

  • 解决在重尾数据和对抗性污染下经典经验风险最小化(ERM)失效的问题,此时标准的集中不等式不再成立。
  • 开发一种鲁棒的ERM框架,在弱矩条件下(如有限二至四阶矩)保持统计效率和高概率性能保证。
  • 确保在重尾或污染数据下,所得估计量的超额风险以快速速率收敛至零。
  • 提供明确依赖于污染水平和样本大小的超额风险理论界,从而在最小分布假设下实现实际的鲁棒学习。

提出的方法

  • 在ERM中用分位数均值估计量替代标准经验均值,仅在有限二阶矩下即可实现次高斯集中。
  • 采用Catoni的鲁棒估计量作为期望的替代代理,提供在弱矩条件下的紧密偏差界。
  • 通过将数据划分为不相交的子样本,对每个子样本计算鲁棒估计量,再取其中位数,以降低对异常值的敏感性。
  • 利用针对鲁棒估计量定制的泛化过程理论和对称化技术,推导出高置信度的超额风险界。
  • 通过涉及梯度估计量比值和余项的稳定性论证,引入一种改进的超额风险控制机制。
  • 将该框架应用于线性回归和逻辑回归,证明其在鲁棒性约束下对标准统计学习问题的适用性。

实验结果

研究问题

  • RQ1在仅具有有限二至四阶矩的条件下,鲁棒经验风险最小化能否实现快速的超额风险收敛速率?
  • RQ2对抗性污染的存在如何影响经典ERM的性能?鲁棒估计量能否缓解这一影响?
  • RQ3分位数均值或Catoni型估计量在多大程度上可以替代ERM中的经验均值,同时保持快速收敛速率?
  • RQ4在鲁棒设置下,超额风险界对污染水平和样本大小的依赖关系如何?
  • RQ5所提出的鲁棒ERM框架能否扩展到线性回归和逻辑回归等实际问题,并提供理论保证?

主要发现

  • 所提出的鲁棒ERM估计量即使在损失函数仅有有限二至四阶矩时,其超额风险界仍能以快速速率收敛至零。
  • 以高概率 $1 - 10e^{-s}$,鲁棒估计量 $ ilde{f}_N$ 的超额风险被限制在 $ ilde{ ho} + C( ho)ig(D^2 ho( ho, ho) ig)ig( rac{ rak{B}^6( ho, ho)}{M_ ho^4 n^2} + rac{s + ho}{N}ig)$ 内,其中 $ ho$ 为污染水平。
  • 超额风险界仅通过项 $ rac{s + ho}{N}$ 以次优方式依赖于污染水平,表明对重尾分布和对抗性异常值均具有鲁棒性。
  • 该框架确保鲁棒估计量 $ ilde{f}_N$ 满足 $ ilde{ ho} \to 0$ 当 $N \to \infty$,即使真实分布具有重尾特性。
  • 在高概率事件 $\mathcal{E}_1$ 下,候选估计量集合 $ {F}( ilde{ ho})$ 位于最优风险的 $7\tilde{ ho}$-邻域内。
  • 理论分析证实,该鲁棒估计量在最小矩假设下仍保持快速收敛速率,且明确依赖于包络范数 $\frak{B}(\ell,\mathcal{F})$ 和鲁棒性参数 $M_\Delta$。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。