[论文解读] On the bias, risk and consistency of sample means in multi-armed bandits
本文对自适应采样、停止、选择和回溯条件下多臂赌博机(MAB)设置中样本均值的偏差、风险和一致性进行了全面分析。识别出四种选择偏差来源,引入有效样本量以界定风险,并利用变分表示和鞅集中不等式,在一般矩条件之下建立了的一致性和风险控制。
The sample mean is among the most well studied estimators in statistics, having many desirable properties such as unbiasedness and consistency. However, when analyzing data collected using a multi-armed bandit (MAB) experiment, the sample mean is biased and much remains to be understood about its properties. For example, when is it consistent, how large is its bias, and can we bound its mean squared error? This paper delivers a thorough and systematic treatment of the bias, risk and consistency of MAB sample means. Specifically, we identify four distinct sources of selection bias (sampling, stopping, choosing and rewinding) and analyze them both separately and together. We further demonstrate that a new notion of \emph{effective sample size} can be used to bound the risk of the sample mean under suitable loss functions. We present several carefully designed examples to provide intuition on the different sources of selection bias we study. Our treatment is nonparametric and algorithm-agnostic, meaning that it is not tied to a specific algorithm or goal. In a nutshell, our proofs combine variational representations of information-theoretic divergences with new martingale concentration inequalities.
研究动机与目标
- 系统分析自适应数据收集下多臂赌博机中样本均值的偏差、风险和一致性。
- 识别并表征四种不同的选择偏差来源:采样、停止、选择和回溯。
- 在完全自适应设置(所有四个自适应组件)下,建立样本均值一致性的充分条件。
- 通过新颖的有效样本量概念,在一般矩条件和尾部条件下推导出样本均值的紧致风险界。
- 提供一种非参数、与算法无关的框架,适用于多种MAB应用场景,无需假设特定算法或目标。
提出的方法
- 使用信息论分歧的变分表示分析偏差和风险。
- 应用新的鞅集中不等式,控制自适应停止和采样下样本均值的行为。
- 引入一种新颖的有效样本量概念,以在一般损失函数下界定样本均值的均方误差。
- 通过Donsker-Varadhan表示和自归一化过程不等式推导风险界,特别针对次高斯臂。
- 采用非参数、与算法无关的方法,不依赖于参数假设或特定的赌博机算法。
- 通过精心设计的示例验证理论结果,说明每种偏差来源的影响。
实验结果
研究问题
- RQ1在涉及自适应采样、停止、选择和回溯的完全自适应MAB设置中,样本均值在何种条件下是一致的?
- RQ2MAB数据收集中的四种不同的选择偏差来源是什么,它们如何相互作用?
- RQ3在自适应设置下,一般矩条件和尾部条件下样本均值的风险如何界定?
- RQ4是否可以使用一种新的有效样本量概念来控制自适应赌博机实验中样本均值的均方误差?
- RQ5与现有针对次高斯臂的风险界相比,所提出的边界在紧致性和普适性方面表现如何?
主要发现
- 在自适应MAB设置中,样本均值存在偏差,本文识别出四种不同的选择偏差来源:采样、停止、选择和回溯。
- 本文证明,在自适应算法满足一般单调性条件时,即使同时存在全部四个自适应组件,样本均值仍具有一致性。
- 引入一种新的有效样本量度量,并证明其在一般损失函数下可对样本均值的均方误差提供紧致界。
- 对于次高斯臂,本文通过自归一化过程推导出风险界,其结果与主界仅相差一个常数因子,且在样本量变异性较高时控制更紧密。
- 所提出的风硏界适用于任意随机时间,而某些现有边界仅限于停止时间。
- 在高变异性设置下,有效样本量归一化优于其他归一化方法(如$ ilde{N}^{ ext{E}}$),能实现更紧致的风险控制。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。