[论文解读] Fast learning rates with heavy-tailed losses
本文通过引入两个关键条件——$L^r$-可积的包络函数与多尺度伯恩斯坦条件——建立了在重尾损失下的经验风险最小化的快速学习速率。在这些假设下,证明了学习速率可任意接近$O(n^{-1})$,并在重尾数据下的$k$-均值聚类中得到验证。
We study fast learning rates when the losses are not necessarily bounded and may have a distribution with heavy tails. To enable such analyses, we introduce two new conditions: (i) the envelope function $\sup_{f \in \mathcal{F}}|\ell \circ f|$, where $\ell$ is the loss function and $\mathcal{F}$ is the hypothesis class, exists and is $L^r$-integrable, and (ii) $\ell$ satisfies the multi-scale Bernstein's condition on $\mathcal{F}$. Under these assumptions, we prove that learning rate faster than $O(n^{-1/2})$ can be obtained and, depending on $r$ and the multi-scale Bernstein's powers, can be arbitrarily close to $O(n^{-1})$. We then verify these assumptions and derive fast learning rates for the problem of vector quantization by $k$-means clustering with heavy-tailed distributions. The analyses enable us to obtain novel learning rates that extend and complement existing results in the literature from both theoretical and practical viewpoints.
研究动机与目标
- 解决损失无界且具有重尾时快速学习速率缺乏理论理解的问题。
- 将现有快速速率理论从有界或次高斯损失扩展至具有多项式或更重尾部的分布。
- 提供可验证的条件,使得在无界损失设定下可实现接近$O(n^{-1})$的快速收敛速率。
- 在具有重尾源分布的$k$-均值聚类上验证所提出的框架。
提出的方法
- 引入包络函数$\sup_{f \in \mathcal{F}} |\ell \circ f|$的$L^r$-可积性($r \geq 2$),以控制无界损失的尾部行为。
- 提出多尺度伯恩斯坦条件作为标准伯恩斯坦条件在无界损失下的推广,支持快速速率分析。
- 在$L^r$-可积性条件下,使用Lederer和van de Geer(2014)的泛函过程上确界浓度不等式,以界 excess risk。
- 建立风险函数的黑塞矩阵与多尺度伯恩斯坦条件之间的联系,确保在最优解附近实现快速收敛。
- 通过在损失函数与假设类在矩和正则性假设下验证条件,将该框架应用于$k$-均值聚类。
- 利用覆盖数与熵条件推导有限样本界,将速率与数据分布的矩数$r$联系起来。
实验结果
研究问题
- RQ1当损失函数具有重尾且无界时,能否实现快速学习速率?
- RQ2在重尾损失存在的情况下,假设类与损失函数的何种条件可实现快速速率?
- RQ3多尺度伯恩斯坦条件是否为无界损失设定下快速速率的充分且可验证条件?
- RQ4所提出的框架能否在重尾数据下的$k$-均值聚类中实现比现有结果更快的收敛速率?
- RQ5数据分布的矩数$r$如何影响可实现的学习速率?
主要发现
- 当包络函数满足$L^r$-可积性($r \geq 2$)时,可实现快于$O(n^{-1/2})$的学习速率。
- 在多尺度伯恩斯坦条件成立下,当$r \to \infty$时,学习速率可任意接近$O(n^{-1})$。
- 对于$k$-均值聚类,若数据的矩数高达$r \geq 4k(d+1)$,则过剩风险以$O(n^{-\beta})$的速率收敛,其中$\beta < \frac{r-1}{r}(1 - 2\sqrt{k(d+1)/r})$。
- 当$r \to \infty$时,收敛速率趋近于$O(n^{-1})$,表明在重尾分布下实现近乎最优性能。
- 结果扩展了Antos等人(2005)与Levrard(2013)的工作,适用于无界支撑,并在类似假设下优于Telgarsky与Dasgupta(2013)的结果。
- 多尺度伯恩斯坦条件被证明是无界损失设定下一类自然且可验证的假设,能够区分微观与宏观尺度下的行为差异。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。