QUICK REVIEW
[论文解读] Bounding the expectation of the supremum of an empirical process over a (weak) vc-major class
Yannick Baraud|arXiv (Cornell University)|Nov 20, 2014
Machine Learning and Algorithms参考文献 5被引用 17
一句话总结
本文为 (弱) VC-major 或 VC-subgraph 类函数上的经验过程的上确界期望提供了非渐近上界,包含显式常数,并相较于通用熵界实现了更优的收敛速率。该方法利用熵函数的凹包络实现局部熵控制,在方差较小时获得更紧的界,尤其在全局熵界过于保守的情况下表现更优。
ABSTRACT
Given a bounded class of functions G and independent random variables X1, . . . , Xn, we provide an upper bound for the expectation of the supremum of the empirical process over elements of G having a small variance. Our bound applies in the cases where G is a VC-subgraph or a VC-major class and it is of smaller order than those one could get by using a universal entropy bound over the whole class G . It also involves explicit constants and does not require the knowledge of the entropy of G
研究动机与目标
- 为 (弱) VC-major 或 VC-subgraph 类上的经验过程的上确界期望开发更紧的上界。
- 改进在类在方差或 L2-范数上局部较小时过于保守的通用熵界。
- 提供具有显式数值常数的非渐近界,避免依赖全局熵的计算。
- 处理函数类被限制在某个给定函数附近(例如在 L2-球内)的情形,此时全局熵界无法捕捉局部复杂性。
提出的方法
- 该方法使用熵函数的凹包络 $\overline{H}(u)$,在不同尺度上控制类的熵。
- 应用对称化与链式技巧,通过 Rademacher 平均界 bounds 经验过程的上确界期望。
- 关键不等式通过在 $[0,1]$ 上对 $\overline{H}$ 的积分来界 bounds $\mathbb{E}[\sup_{f \in \mathcal{F}} |\sum_{i=1}^n \varepsilon_i f(X_i)|]$。
- 该方法利用 VC-major 性质,确保水平集族 $\{f > u\}$ 是 VC 的,从而实现熵控制。
- 通过变换使函数中心化,将问题转化为具有可控方差的有界类。
- 证明依赖于指示函数的序列闭包论证,以及 VC 类的性质,以保持维数界。
实验结果
研究问题
- RQ1我们能否在通用熵界之外,为 VC-major 类上的经验过程上确界期望获得更紧的界?
- RQ2如何利用局部方差或 L2-范数控制来改进经验过程界中的收敛速率?
- RQ3在给定函数(如常数函数或快速递减函数)附近的熵结构在决定局部类的复杂性中起什么作用?
- RQ4我们能否推导出非渐近界,包含显式常数,而无需计算整个函数类的全局熵?
主要发现
- 对 $\mathbb{E}[Z(\mathcal{F})]$ 的界为 $\sqrt{n} \cdot \overline{H}(\sigma \vee a - \sigma \log(\sigma \vee a))$ 阶,当 $\sigma$ 较小时严格小于通用熵界。
- 对于类 $\mathcal{F} = \mathcal{G} \cap \mathcal{B}(g_0, r)$,其中 $\mathcal{G}$ 为 VC-subgraph 或 VC-major 类,当 $g_0$ 为常数时,界显著改善,因为局部熵远小于全局熵。
- 该方法得到的界为 $\sqrt{n\sigma} + \sigma^{-1}$ 阶,但常数更小,尤其在 $\sigma \ll 1$ 时优势明显。
- 该界是非渐近的,适用于独立但未必同分布的随机变量。
- 证明表明,若 $\mathcal{F}$ 是 VC-major 且维数为 $d$,则 $\mathcal{G} = \{g_f = \frac{1}{2}(f - \mathbb{E}[f(X_1)])\}$ 是弱 VC-major,且维数至多为 $d$,从而保持了复杂性结构。
- 该结果改进了 Giné 和 Koltchinski (2006) 的例 3.8,其中已知期望小于 $C(\sqrt{n\sigma} + \sigma^{-1})$,当 $g_0 = 0$ 时。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。