Skip to main content
QUICK REVIEW

[论文解读] New Potential-Based Bounds for Prediction with Expert Advice

Vladimir A. Kobzar, Robert V. Kohn|arXiv (Cornell University)|Nov 5, 2019
Advanced Bandit Algorithms Research参考文献 24被引用 4
一句话总结

本文通过利用最优控制理论导出的偏微分方程(PDE)的次解与超解,提出了在线预测中专家建议问题的新潜在函数界。该方法为任意数量专家的有限时域博弈提供了更紧致的遗憾界,对两名和三名专家实现了最优的一阶项,并通过PDE的闭式解在特定参数区域内超越了先前的最先进结果。

ABSTRACT

This work addresses the classic machine learning problem of online prediction with expert advice. We consider the finite-horizon version of this zero-sum, two-person game. Using verification arguments from optimal control theory, we view the task of finding better lower and upper bounds on the value of the game (regret) as the problem of finding better sub- and supersolutions of certain partial differential equations (PDEs). These sub- and supersolutions serve as the potentials for player and adversary strategies, which lead to the corresponding bounds. To get explicit bounds, we use closed-form solutions of specific PDEs. Our bounds hold for any given number of experts and horizon; in certain regimes (which we identify) they improve upon the previous state of the art. For two and three experts, our bounds provide the optimal leading order term.

研究动机与目标

  • 通过潜在函数开发有限时域专家问题的更紧致、非渐近的遗憾界。
  • 将潜在函数方法扩展至下界,通过PDE的次解与超解推导下界。
  • 识别新界优于现有最先进结果的参数区域。
  • 利用特定PDE的解,为任意数量专家和时域提供显式、闭式遗憾界。
  • 统一并推广先前基于PDE的在线学习最优策略方法。

提出的方法

  • 将专家问题建模为零和博弈,并利用最优控制中的验证论证,推导PDE的次解与超解作为潜在函数。
  • 将博弈的值函数视为哈密顿-雅可比-贝尔曼PDE的解,并通过次解与超解构造界。
  • 应用线性和非线性PDE(如热方程及其非线性类比)的闭式解,生成显式界。
  • 通过最大潜在函数构造上界,通过对称潜在函数构造下界,并利用高阶PDE导数控制误差项。
  • 使用数值积分和随机游走近似方法计算并验证界,以供比较。
  • 证明对于N=2和N=3名专家,该界在首项系数上是紧致的。

实验结果

研究问题

  • RQ1能否系统性地利用PDE的次解与超解,推导专家问题中遗憾的上界与下界?
  • RQ2哪些类别的潜在函数能为一般N和有限T提供更紧致的非渐近遗憾界?
  • RQ3基于PDE的界与已知的渐近与非渐近结果相比如何,特别是在小N情况下?
  • RQ4PDE的闭式解能否导出显式、可计算的遗憾界,并优于先前工作?
  • RQ5在哪些参数区域内,新界对N=2和N=3实现了最优的一阶项?

主要发现

  • 对于两名专家,新界在遗憾中实现了最优的一阶项,与已知的非渐近最优策略一致。
  • 对于三名专家,该界同样实现了最优的一阶项,优于先前的最先进结果。
  • 该方法为任意数量专家和有限时域T提供了显式、闭式遗憾界,其来源于特定PDE的解。
  • 在适当的PDE参数化下,上界误差项为O(N log |t|),优于先前的潜在函数方法。
  • 通过对称潜在函数导出的下界与独立同分布高斯情形下的已知下界一致,确认了在渐近极限下的紧致性。
  • 该框架统一并扩展了先前基于PDE的方法,通过次解与超解实现上界与下界。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。