[论文解读] Distribution oblivious, risk-aware algorithms for multi-armed bandits with unbounded rewards
本文提出了适用于无界奖励的多臂赌博机中最佳臂识别的分布无关、风险感知算法,通过优化期望奖励与条件风险价值(CVaR)之间的权衡。提出了一种新颖的CVaR估计器及重尾分布的集中不等式,实现了无需事先了解奖励分布矩或次优性差距的可证明误差界。
Classical multi-armed bandit problems use the expected value of an arm as a metric to evaluate its goodness. However, the expected value is a risk-neutral metric. In many applications like finance, one is interested in balancing the expected return of an arm (or portfolio) with the risk associated with that return. In this paper, we consider the problem of selecting the arm that optimizes a linear combination of the expected reward and the associated Conditional Value at Risk (CVaR) in a fixed budget best-arm identification framework. We allow the reward distributions to be unbounded or even heavy-tailed. For this problem, our goal is to devise algorithms that are entirely distribution oblivious, i.e., the algorithm is not aware of any information on the reward distributions, including bounds on the moments/tails, or the suboptimality gaps across arms. In this paper, we provide a class of such algorithms with provable upper bounds on the probability of incorrect identification. In the process, we develop a novel estimator for the CVaR of unbounded (including heavy-tailed) random variables and prove a concentration inequality for the same, which could be of independent interest. We also compare the error bounds for our distribution oblivious algorithms with those corresponding to standard non-oblivious algorithms. Finally, numerical experiments reveal that our algorithms perform competitively when compared with non-oblivious algorithms, suggesting that distribution obliviousness can be realised in practice without incurring a significant loss of performance.
研究动机与目标
- 为解决经典多臂赌博机算法在风险管理至关重要的应用中依赖风险中性期望值度量的局限性。
- 设计完全分布无关的算法——无需事先了解奖励分布的矩、尾部行为或次优性差距——同时仍能保证强性能保障。
- 在固定预算的最佳臂识别框架下,优化期望奖励与条件风险价值(CVaR)的线性组合。
- 在无界或重尾奖励分布下,建立错误臂识别概率的可证明上界。
提出的方法
- 设计一种针对无界和重尾随机变量的CVaR新颖估计器,使风险感知决策无需分布假设成为可能。
- 证明所提出的CVaR估计器的集中不等式,这对于推导有限样本性能保证至关重要。
- 构建一类分布无关的算法,基于估计的CVaR和期望奖励平衡探索与利用,而无需使用任何分布信息。
- 采用固定预算框架,确保即使在奖励分布未知且可能为重尾的情况下,算法也能以高概率选择最佳臂。
- 将CVaR估计器集成到赌博机学习框架中,根据风险调整后的性能估计动态分配对各臂的尝试次数。
- 推导错误识别概率的理论上的上界,这些上界仅依赖于问题的固有难度,而不依赖于分布参数。
实验结果
研究问题
- RQ1我们能否设计出适用于无界奖励分布的分布无关算法,以有效平衡期望奖励与CVaR所衡量的风险?
- RQ2当奖励分布无界或为重尾且无任何分布信息时,CVaR的稳健且可证明准确的估计器是什么?
- RQ3与利用矩或尾部行为知识的非无关算法相比,分布无关算法的误差界如何?
- RQ4尽管缺乏分布假设,分布无关算法在实践中能否实现具有竞争力的性能?
- RQ5所提出的CVaR估计器在一般无界分布下表现出怎样的集中性质?
主要发现
- 所提出的CVaR估计器是首个在无需了解矩或尾部衰减速率的前提下,为无界且可能重尾的分布提供可证明集中界估计器。
- 本文建立了错误识别概率的理论可证明上界,且该上界在仅假设一阶矩存在的情况下即成立,无需对奖励分布作其他假设。
- 数值实验表明,分布无关算法与非无关算法相比表现具有竞争力,表明分布无知不会带来显著的性能损失。
- 所提出的CVaR估计器的集中不等式本身具有独立兴趣,可能在赌博机设置之外也具有适用性。
- 所提算法的误差界在数量级上与非无关算法相当,证明了在不牺牲理论保证的前提下,实现分布无关性是可行的。
- 该框架成功将风险感知(通过CVaR)整合到最佳臂识别问题中,同时保持了分布无关性,这是赌博机文献中的一个新颖贡献。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。