Skip to main content
QUICK REVIEW

[论文解读] Online Learning: Stochastic and Constrained Adversaries

Alexander Rakhlin, Karthik Sridharan|arXiv (Cornell University)|Apr 27, 2011
Advanced Bandit Algorithms Research参考文献 19被引用 12
一句话总结

本文提出了一种统一的极小极大框架,用于在混合数据假设下的在线学习,介于独立同分布(i.i.d.)随机设置与最坏情况对抗设置之间。通过利用序列对称化和与分布相关的 Rademacher 复杂度,推导出变体类型的遗憾界,并在 i.i.d. 对抗者下建立了在线学习与批量学习可学习性的等价性,同时表明即使在 Littlestone 维度无限的情况下,半空间在平滑噪声下也变得可学习。

ABSTRACT

Learning theory has largely focused on two main learning scenarios. The first is the classical statistical setting where instances are drawn i.i.d. from a fixed distribution and the second scenario is the online learning, completely adversarial scenario where adversary at every time step picks the worst instance to provide the learner with. It can be argued that in the real world neither of these assumptions are reasonable. It is therefore important to study problems with a range of assumptions on data. Unfortunately, theoretical results in this area are scarce, possibly due to absence of general tools for analysis. Focusing on the regret formulation, we define the minimax value of a game where the adversary is restricted in his moves. The framework captures stochastic and non-stochastic assumptions on data. Building on the sequential symmetrization approach, we define a notion of distribution-dependent Rademacher complexity for the spectrum of problems ranging from i.i.d. to worst-case. The bounds let us immediately deduce variation-type bounds. We then consider the i.i.d. adversary and show equivalence of online and batch learnability. In the supervised setting, we consider various hybrid assumptions on the way that x and y variables are chosen. Finally, we consider smoothed learning problems and show that half-spaces are online learnable in the smoothed model. In fact, exponentially small noise added to adversary's decisions turns this problem with infinite Littlestone's dimension into a learnable problem.

研究动机与目标

  • 通过分析介于 i.i.i.d. 统计学习与最坏情况在线学习之间的中间数据假设,弥合两者之间的差距。
  • 为在受限对抗者(包括随机和光滑模型)下的遗憾最小化,开发一个通用的理论框架。
  • 在对抗者生成 i.i.i.d. 数据时,建立在线学习与批量学习可学习性的等价性。
  • 分析输入和输出变量在不同假设下生成的混合模型。
  • 证明即使 Littlestone 维度为无穷大,半空间在平滑噪声下也变得可学习。

提出的方法

  • 提出一种在线学习的极小极大公式化,其中对抗者被限制在特定的动作集合中,从而能够分析从 i.i.i.d. 到最坏情况的整个谱系。
  • 引入一种针对受限对抗者的序列博弈的、与分布相关的 Rademacher 复杂度,通过序列对称化推导得出。
  • 通过使用 i.i.d. Rademacher 标签的随机过程构造,对博弈的极小极大值进行下界估计,将其与 Rademacher 混合过程的期望上确界联系起来。
  • 通过分析条件分布和标签依赖结构,将该框架应用于推导变体类型的遗憾界。
  • 通过证明当且仅当函数类具有有限的组合维数时,极小极大遗憾收敛于零,从而在 i.i.i.d. 数据下建立在线学习与批量学习可学习性的等价性。
  • 证明在添加指数级小噪声到对抗决策(即平滑模型)时,即使 Littlestone 维度为无穷大,半空间也变得可在线学习。

实验结果

研究问题

  • RQ1能否开发一个统一的框架,用于分析介于 i.i.i.d. 和最坏情况之间的中间数据假设下的在线学习?
  • RQ2当对抗者生成 i.i.i.d. 数据时,在线学习与批量学习可学习性之间存在何种关系?
  • RQ3在输入和输出在不同假设下生成的混合模型中,极小极大遗憾如何表现?
  • RQ4Rademacher 复杂度的概念能否扩展到具有受限对抗者的序列博弈?
  • RQ5在添加小噪声(即平滑模型)时,是否能使原本不可学习的问题(如 Littlestone 维度为无穷大的半空间)在在线设置下变得可学习?

主要发现

  • 当对抗者生成 i.i.i.d. 数据时,极小极大遗憾收敛于零,当且仅当函数类具有有限的组合维数,这与经典统计学习结果一致。
  • 为整个数据假设谱系定义了一种与分布相关的 Rademacher 复杂度,从而能够推导出变体类型的遗憾界。
  • 本文在 i.i.i.d. 数据下建立了在线学习与批量学习可学习性的等价性,表明两种设置具有相同的极小极大遗憾率。
  • 在平滑模型下,即对抗者添加指数级小噪声时,即使 Littlestone 维度为无穷大,半空间也变得可在线学习。
  • 博弈的极小极大值由函数类上 Rademacher 过程的期望上确界下界界定,该下界在 i.i.i.d. 标签假设下被证明等于与分布相关的 Rademacher 复杂度。
  • 该框架成功地通过单一极小极大分析统一了最坏情况、i.i.i.d. 和平滑模型,为分析具有结构化对抗者的序列决策问题提供了一个通用工具。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。