Skip to main content
QUICK REVIEW

[论文解读] Following the Leader and Fast Rates in Linear Prediction: Curved Constraint Sets and Other Regularities

Ruitong Huang, Tor Lattimore|arXiv (Cornell University)|Feb 10, 2017
Advanced Bandit Algorithms Research被引用 13
一句话总结

本文证明了在在线线性预测中,即使损失函数不具曲率(例如非强凸),只要约束集边界具有曲率——具体而言,当决策集为强凸且损失向量的均值范数远离零时——Follow the Leader (FTL) 算法仍能实现对数 regret。这揭示了一种新的在线学习中快速收敛率的机制,该机制独立于损失函数的曲率。

ABSTRACT

The follow the leader (FTL) algorithm, perhaps the simplest of all online learning algorithms, is known to perform well when the loss functions it is used on are convex and positively curved. In this paper we ask whether there are other "lucky" settings when FTL achieves sublinear, "small" regret. In particular, we study the fundamental problem of linear prediction over a non-empty convex, compact domain. Amongst other results, we prove that the curvature of the boundary of the domain can act as if the losses were curved: In this case, we prove that as long as the mean of the loss vectors have positive lengths bounded away from zero, FTL enjoys a logarithmic growth rate of regret, while, e.g., for polytope domains and stochastic data it enjoys finite expected regret. Building on a previously known meta-algorithm, we also get an algorithm that simultaneously enjoys the worst-case guarantees and the bound available for FTL.

研究动机与目标

  • 研究 FTL 是否能在损失函数无曲率的情况下,于在线线性预测中实现次线性 regret。
  • 识别决策集或数据中可导致 FTL 实现快速 regret 率的结构性规律。
  • 证明:当约束集边界的强凸性与非退化的损失均值相结合时,可在对抗性设置下实现对数 regret。
  • 提供一个匹配的下界,以证明所推导的 regret 上界是紧致的。

提出的方法

  • 作者在凸且紧致的决策集上分析 FTL 算法在在线线性预测中的表现。
  • 引入一个几何条件:约束集边界的强凸性(曲率),该条件作为正则性,可实现快速收敛率。
  • 通过利用决策集边界的曲率来推导 regret 上界,即使损失函数为线性亦成立。
  • 构建一个元算法,结合最坏情况下的鲁棒性与在有利正则性条件下的快速收敛率。
  • 关键技术工具包括集中不等式以及对偶空间的几何分析,特别是法锥与支撑函数的相关分析。
  • 证明技术通过损失向量与约束集边界法线之间夹角的有界性来控制 regret,利用曲率确保充分进展。

实验结果

研究问题

  • RQ1当损失函数为线性但约束集具有曲率边界时,FTL 是否能在在线线性预测中实现次线性 regret?
  • RQ2决策集边界的强凸性是否可作为正则性,使 FTL 实现快速 regret 率,即使损失函数无曲率?
  • RQ3在上述几何正则性条件下,FTL 可实现的最优 regret 率是多少?该上界是否紧致?
  • RQ4是否存在单一算法,既能保证最坏情况下的鲁棒性,又能在有利条件下实现快速收敛率?

主要发现

  • 当决策集为强凸且损失向量的均值范数远离零时,FTL 在在线线性预测中可实现对数 regret。
  • 约束集边界的曲率可作为损失曲率的替代机制,即使损失为线性,也能实现快速收敛率。
  • 对于具有随机数据的多面体域,FTL 实现有限期望 regret,表明其相比最坏情况下的线性 regret 有显著改进。
  • 提出一种元算法,结合最坏情况保证与在有利条件下的 FTL 快速收敛率。
  • 本文建立了匹配的下界,表明所推导的对数 regret 上界在对数因子范围内是紧致的。
  • 结果表明,约束集中的几何正则性可与损失函数的曲率一样,有效促进在线学习中的快速收敛率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。