[论文解读] No fast exponential deviation inequalities for the progressive mixture rule
本文表明,尽管渐进混合规则在实现最优 $ O(1/n) $ 期望风险率方面表现良好,但其并不具备快速的指数型偏差不等式——其偏差界仅为 $ O(1/\sqrt{n}) $,即使在标准假设(如指数凸性和对称损失)下亦然。这表明其高概率性能表现次优,相较于基于期望的保证存在差距。
We consider the learning task consisting in predicting as well as the best function in a finite reference set G up to the smallest possible additive term. If R(g) denotes the generalization error of a prediction function g, under reasonable assumptions on the loss function (typically satisfied by the least square loss when the output is bounded), it is known that the progressive mixture rule g_n satisfies E R(g_n) < min_{g in G} R(g) + C (log|G|)/n where n denotes the size of the training set, E denotes the expectation w.r.t. the training set distribution and C denotes a positive constant. This work mainly shows that for any training set size n, there exist a>0, a reference set G and a probability distribution generating the data such that with probability at least a R(g_n) > min_{g in G} R(g) + c sqrt{[log(|G|/a)]/n}, where c is a positive constant. In other words, surprisingly, for appropriate reference set G, the deviation convergence rate of the progressive mixture rule is only of order 1/sqrt{n} while its expectation convergence rate is of order 1/n. The same conclusion holds for the progressive indirect mixture rule. This work also emphasizes on the suboptimality of algorithms based on penalized empirical risk minimization on G.
研究动机与目标
- 探究渐进混合规则是否在已知其具有最优期望风险率的前提下,也能实现快速的指数型偏差不等式。
- 分析渐进混合规则及其间接变体在标准损失假设下的偏差行为。
- 挑战一种常见假设,即良好的期望性能意味着强的高概率界。
- 强调惩罚的经验风险最小化在实现快速偏差速率方面存在次优性。
提出的方法
- 通过构造一个特定的数据生成分布和参考集 $ \mathcal{G} $,以证明偏差概率的下界。
- 利用霍夫丁不等式和矩生成函数技术,对经验风险最小化器的期望风险进行界控。
- 通过优化边差异的超立方体分布族,应用阿苏阿德型下界。
- 将渐进混合规则定义为基于累积损失顺序更新后验权重的时间平均预测。
- 关键不等式涉及利用边差异 $ d_{\textnormal{I}} $ 和归一化偏差项,对在 $ \mathcal{G} $ 中最优函数上的风险超额进行界控。
- 构造过程利用了损失函数的对称性和可容许性,以确保边差异为正,从而实现非平凡的下界。
实验结果
研究问题
- RQ1在标准的指数凸性和对称损失假设下,渐进混合规则能否实现快速的指数型偏差不等式?
- RQ2渐进混合规则的 $ O(1/n) $ 期望风险率是否对应于相应的 $ O(1/n) $ 偏差率?
- RQ3在 $ n $ 和 $ |\mathcal{G}| $ 的意义上,渐进混合规则的最小可能偏差率是多少?
- RQ4为何惩罚的经验风险最小化方法在此设置下无法实现快速偏差速率?
主要发现
- 对于任意 $ n $,存在一个分布和参考集 $ \mathcal{G} $,使得以至少 $ \epsilon > 0 $ 的概率,渐进混合规则的风险超额至少为 $ c\sqrt{\frac{\log(|\mathcal{G}|\epsilon^{-1})}{n}} $。
- 尽管期望速率是 $ O(1/n) $,其偏差收敛速率仅为 $ O(1/\sqrt{n}) $,表明高概率性能存在根本性差距。
- 该次优偏差速率同样适用于渐进间接混合规则。
- 由于“中心处表现良好”和可容许性假设,边差异 $ d_{\textnormal{I}} $ 严格为正,从而可实现非平凡下界。
- 通过设定 $ \tilde{d}_{\textnormal{II}} = \min\left(\frac{\tilde{m}}{4n}, 1\right) $ 优化下界,从而得出所述偏差速率。
- 结果表明,即使期望速率最优,惩罚的经验风险最小化在实现快速偏差速率方面仍属次优。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。