[论文解读] Approximations of the Restless Bandit Problem
本文提出了在 $φ$-mixing 奖励依赖关系下,对非平稳多臂赌博机问题的可 tractable 近似方法,表明经过修改的 UCB 算法在最高平稳均值方面实现了对数 regret。关键贡献在于,其 regret 上界依赖于 $φ$-mixing 系数之和,在依赖关系消失时可恢复 i.i.d. 情况。
The multi-armed restless bandit problem is studied in the case where the pay-off distributions are stationary $φ$-mixing. This version of the problem provides a more realistic model for most real-world applications, but cannot be optimally solved in practice, since it is known to be PSPACE-hard. The objective of this paper is to characterize a sub-class of the problem where {\em good} approximate solutions can be found using tractable approaches. Specifically, it is shown that under some conditions on the $φ$-mixing coefficients, a modified version of UCB can prove effective. The main challenge is that, unlike in the i.i.d. setting, the distributions of the sampled pay-offs may not have the same characteristics as those of the original bandit arms. In particular, the $φ$-mixing property does not necessarily carry over. This is overcome by carefully controlling the effect of a sampling policy on the pay-off distributions. Some of the proof techniques developed in this paper can be more generally used in the context of online sampling under dependence. Proposed algorithms are accompanied with corresponding regret analysis.
研究动机与目标
- 解决在依赖性非 i.i.d. 奖励下,非平稳多臂赌博机问题的计算不可行性。
- 识别出非平稳多臂赌博机问题的一个子类,使得可设计出可 tractable 的近似算法。
- 分析 $φ$-mixing 依赖性对 regret 和最优策略性能的影响。
- 设计一种适用于弱依赖奖励的 UCB 类算法,并建立其 regret 上界。
- 提供关于混合系数和平稳均值的近似质量的理论保证。
提出的方法
- 假设奖励序列是平稳的 $φ$-mixing,以建模长程时间依赖和跨臂依赖。
- 提出一种修改后的 UCB 算法,以考虑采样策略对依赖奖励分布的影响。
- 使用专为 $φ$-mixing 过程设计的集中不等式,以界定所提算法的 regret。
- 通过将期望 regret 分解为涉及平稳均值和依赖结构的分量,建立 regret 上界。
- 对协方差函数施加 Hölder 连续性假设,以控制采样对分布特征的影响。
- 采用条件独立性论证和密度有界技术,推导出次优臂选择对 regret 的贡献上界。
实验结果
研究问题
- RQ1在何种 $φ$-mixing 系数条件下,非平稳多臂赌博机问题中的最优切换策略可由选择具有最高平稳均值的臂来近似?
- RQ2如何将乐观探索策略适配于依赖奖励过程,以实现次线性 regret?
- RQ3regret 上界如何依赖于奖励过程中时间依赖和跨臂依赖的强度?
- RQ4$φ$-mixing 框架能否用于将 UCB 类算法推广至非 i.i.d. 奖励?
- RQ5基于平稳均值的松弛问题的近似误差如何随奖励过程中依赖程度的变化而变化?
主要发现
- 选择具有最高平稳均值的臂所带来的近似误差是有界的,并且随着 $φ_1$ 减小而减小,表明依赖性减弱。
- 在 $φ$-mixing 假设下,修改后的 UCB 算法在最高平稳均值方面实现了对数 regret。
- regret 上界与 $φ$-mixing 系数之和 $\sum_i \varphi_i$ 成比例,在 i.i.i.d. 情况下可恢复 Auer 等人(2002)的经典对数 regret 上界。
- 所开发的证明技术具有通用性,可推广至依赖过程下的在线采样,特别是通过条件独立性和密度有界技术。
- 分析表明,通过仔细处理采样奖励的联合分布,可控制采样策略引入的依赖性。
- 对于 Hölder-连续协方差函数的情况,regret 上界可进一步优化,并显示其依赖于指数 $\alpha$ 和系数 $c$。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。