[论文解读] Adaptive Experimental Design with Temporal Interference: A Maximum Likelihood Approach
本文提出了一种新颖的自适应实验设计,用于在存在时间干扰的系统中比较两种策略,采用非参数最大似然估计(MLE)一致且高效地估计稳态奖励差异。通过利用鞅分析和泊松方程,将问题简化为凸优化,从而实现一种在线、渐近高效的实验设计,最小化处理效应估计的方差。
Suppose an online platform wants to compare a treatment and control policy, e.g., two different matching algorithms in a ridesharing system, or two different inventory management algorithms in an online retail site. Standard randomized controlled trials are typically not feasible, since the goal is to estimate policy performance on the entire system. Instead, the typical current practice involves dynamically alternating between the two policies for fixed lengths of time, and comparing the average performance of each over the intervals in which they were run as an estimate of the treatment effect. However, this approach suffers from *temporal interference*: one algorithm alters the state of the system as seen by the second algorithm, biasing estimates of the treatment effect. Further, the simple non-adaptive nature of such designs implies they are not sample efficient. We develop a benchmark theoretical model in which to study optimal experimental design for this setting. We view testing the two policies as the problem of estimating the steady state difference in reward between two unknown Markov chains (i.e., policies). We assume estimation of the steady state reward for each chain proceeds via nonparametric maximum likelihood, and search for consistent (i.e., asymptotically unbiased) experimental designs that are efficient (i.e., asymptotically minimum variance). Characterizing such designs is equivalent to a Markov decision problem with a minimum variance objective; such problems generally do not admit tractable solutions. Remarkably, in our setting, using a novel application of classical martingale analysis of Markov chains via Poisson's equation, we characterize efficient designs via a succinct convex optimization problem. We use this characterization to propose a consistent, efficient online experimental design that adaptively samples the two Markov chains.
研究动机与目标
- 解决在线策略评估中的时间干扰问题,其中先前策略的运行会偏差性能估计。
- 为估计两个马尔可夫链之间的稳态奖励差异,开发一种一致且渐近高效的实验设计。
- 在非参数最大似然估计下,对最小方差的最优采样策略进行表征。
- 通过识别可处理的凸优化公式,克服方差最小化马尔可夫决策问题的不可解性。
提出的方法
- 将每种策略建模为在共同状态空间上的马尔可夫链,目标是估计稳态奖励差异。
- 使用非参数最大似然估计(MLE)来估计长期平均奖励,而无需事先知道转移概率或奖励。
- 应用鞅分析和泊松方程,推导出高效设计的理论表征。
- 推导出一个凸优化问题,该问题完全表征了在时间平均正则性(TAR)条件下的渐近高效采样策略。
- 构建一种自适应在线设计,根据当前状态和估计的转移概率动态选择运行哪个策略。
- 通过实时求解凸优化问题并使用估计参数,实现一致且方差最小化的策略。
实验结果
研究问题
- RQ1当仅能进行一次系统运行时,如何使实验设计在存在时间干扰的情况下保持一致和高效?
- RQ2在某些条件下,未知基本参数的方差最小化马尔可夫决策问题能否以闭式解求解?
- RQ3在此设置下,最小化MLE渐近方差的高效采样策略的理论表征是什么?
- RQ4泊松方程和鞅分析的新应用是否能产生适用于自适应设计的可处理优化问题?
- RQ5所得到的设计能否以一致性与效率保证在线实现?
主要发现
- 所提出的自适应设计在所有时间平均正则(TAR)策略中实现了渐近最小方差。
- 尽管方差最小化MDP本身本质上难以求解,但高效设计的表征简化为一个简洁的凸优化问题。
- 在TAR条件下,非参数MLE是一致的,确保收敛到真实的稳态奖励差异。
- MLE的渐近方差通过鞅差正交性和大数定律论证,收敛到所有TAR策略中可能的最小值。
- 在线设计以概率收敛到真实处理效应,状态-动作对的经验频率收敛速率受$O(1/n)$界约束。
- 该解在模型误设下具有鲁棒性,即在弱正则性条件(TAR)下,即使不假设链的参数形式,一致性与效率依然成立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。