[论文解读] Approachability in unknown games: Online learning meets multi-objective optimization
本文提出了一种新颖的框架,用于在未知博弈中实现可接近性,其中决策者在事先不了解博弈结构的情况下,观察任意向量值奖励。该方法提出了一种策略,基于事后观察到的收益自适应地趋近于一个最小且可实现的目标集,避免了投影操作,并通过在各轮次中利用标量后悔最小化,优于凸包松弛方法。
In the standard setting of approachability there are two players and a target set. The players play repeatedly a known vector-valued game where the first player wants to have the average vector-valued payoff converge to the target set which the other player tries to exclude it from this set. We revisit this setting in the spirit of online learning and do not assume that the first player knows the game structure: she receives an arbitrary vector-valued reward vector at every round. She wishes to approach the smallest ("best") possible set given the observed average payoffs in hindsight. This extension of the standard setting has implications even when the original target set is not approachable and when it is not obvious which expansion of it should be approached instead. We show that it is impossible, in general, to approach the best target set in hindsight and propose achievable though ambitious alternative goals. We further propose a concrete strategy to approach these goals. Our method does not require projection onto a target set and amounts to switching between scalar regret minimization algorithms that are performed in episodes. Applications to global cost minimization and to approachability under sample path constraints are considered.
研究动机与目标
- 开发一种在博弈结构事先未知、仅能观测到向量值奖励的未知博弈中可接近性的理论。
- 在原始目标集不可接近时,定义一个事后有意义且可实现的目标集。
- 通过引入一个更具雄心但依然可行的替代目标,克服标准可接近性与凸包松弛方法的局限性。
- 设计一种无需将结果投影到目标集上的实用算法,提升计算可行性。
- 将可接近性理论扩展至全局成本最小化与样本路径约束等问题,而无需假设已知博弈动态。
提出的方法
- 该方法将问题表述为在无收益函数先验知识的向量值老虎机设置中最小化后悔。
- 提出一种策略,在各轮次中切换使用标量后悔最小化算法,以适应观测到的平均收益。
- 通过直接优化基于事后个体响应推导出的目标集,而非投影到固定集合,从而避免投影操作。
- 基于个体响应函数的凸化,定义了一类新的目标集,证明其可实现性。
- 利用对手动作的经验频率来估计最优响应,并根据观测数据动态调整目标函数。
- 通过将约束集视为目标函数估计的一部分,该框架可推广至约束优化问题。
实验结果
研究问题
- RQ1当博弈结构未知且仅能观测到向量值奖励时,决策者是否能在事后趋近最小可能的目标集?
- RQ2在未知博弈中,是否可能实现一个严格小于最优响应函数凸包的目标集?
- RQ3当最优事后目标集不可达时,可以实现哪些替代目标?
- RQ4在高维或复杂场景下,如何在不需投影到目标集的情况下实现可接近性?
- RQ5该框架是否可应用于全局成本最小化与样本路径约束问题,而无需假设已知博弈动态?
主要发现
- 在一般情况下,即使是最理想的目标集(即事后最优目标集)也难以实现。
- 虽然最优响应函数的凸包是可实现的,但其目标过于保守,未能充分捕捉自适应学习的全部潜力。
- 通过提出的新一类可实现且更具雄心的目标集(基于个体响应函数的凸化),可利用所提策略实现趋近。
- 所提算法避免了投影操作,因此在计算效率上优于标准可接近性方案。
- 该方法推广了现有在约束后悔最小化与全局成本最小化中的结果,为未知博弈提供了一个统一的理论框架。
- 在标量收益与约束的特殊情况下,该方法恢复了已知结果,但未实现改进,验证了其与既有工作的兼容性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。