[論文レビュー] Adaptive Experimental Design with Temporal Interference: A Maximum Likelihood Approach
本稿は、時系列干渉が生じるシステムにおける2つの方策の比較に向け、非パラメトリック最尤推定を用いて定常状態報酬差を一貫的かつ効率的に推定する、新しい適応的実験設計を提案する。マルティングール解析とポアソン方程式を活用することで、問題を凸最適化に還元し、オンラインで実行可能かつ漸近的に効率的な設計を実現し、処置効果推定の分散を最小化する。
Suppose an online platform wants to compare a treatment and control policy, e.g., two different matching algorithms in a ridesharing system, or two different inventory management algorithms in an online retail site. Standard randomized controlled trials are typically not feasible, since the goal is to estimate policy performance on the entire system. Instead, the typical current practice involves dynamically alternating between the two policies for fixed lengths of time, and comparing the average performance of each over the intervals in which they were run as an estimate of the treatment effect. However, this approach suffers from *temporal interference*: one algorithm alters the state of the system as seen by the second algorithm, biasing estimates of the treatment effect. Further, the simple non-adaptive nature of such designs implies they are not sample efficient. We develop a benchmark theoretical model in which to study optimal experimental design for this setting. We view testing the two policies as the problem of estimating the steady state difference in reward between two unknown Markov chains (i.e., policies). We assume estimation of the steady state reward for each chain proceeds via nonparametric maximum likelihood, and search for consistent (i.e., asymptotically unbiased) experimental designs that are efficient (i.e., asymptotically minimum variance). Characterizing such designs is equivalent to a Markov decision problem with a minimum variance objective; such problems generally do not admit tractable solutions. Remarkably, in our setting, using a novel application of classical martingale analysis of Markov chains via Poisson's equation, we characterize efficient designs via a succinct convex optimization problem. We use this characterization to propose a consistent, efficient online experimental design that adaptively samples the two Markov chains.
研究の動機と目的
- オンライン方策評価における時系列干渉に対処し、先行方策の実行が性能推定をバイアスすることを防ぐ。
- 2つのマルコフ連鎖間の定常状態報酬差を推定する一貫的かつ漸近的に効率的な実験設計を開発する。
- 最小分散を達成する非パラメトリック最尤推定における最適なサンプリング戦略を特定する。
- 分散最小化マルコフ決定問題の扱いにくさを克服し、解が得られる凸最適化形式を同定する。
提案手法
- 各々の方策を共通の状態空間上のマルコフ連鎖としてモデル化し、定常状態報酬差を推定することを目的とする。
- 遷移確率や報酬に関する事前知識がなくとも、長期間平均報酬を推定できる非パラメトリック最尤推定(MLE)を用いる。
- マルティングール解析とポアソン方程式を適用し、効率的設計の理論的特徴付けを導出する。
- 時間平均正則性(TAR)の下で、漸近的に効率的なサンプリング方策を完全に特徴付ける凸最適化問題を導出する。
- 現在の状態と推定された遷移確率に基づいて、どの方策を実行するかを動的に選択する適応的オンライン設計を構築する。
- 推定されたパラメータを用いてリアルタイムで凸最適化問題を解くことで、一貫的かつ分散最小化方策を実装する。
実験結果
リサーチクエスチョン
- RQ11回のシステム実行しか得られない状況において、時系列干渉が存在する中で、実験設計を一貫的かつ効率的に行う方法は何か?
- RQ2未知の基本的要因を有する分散最小化マルコフ決定問題が、特定の条件下で閉形式で解けるか?
- RQ3この設定において、MLEの漸近的分散を最小化する効率的サンプリング方策の理論的特徴付けは何か?
- RQ4ポアソン方程式とマルティングール解析の新規応用により、適応的設計のための取り扱いやすい最適化問題が得られるか?
- RQ5得られた設計は、一貫性と効率性の保証のもとでオンラインで実装可能か?
主な発見
- 提案された適応的設計は、すべての時間平均正則(TAR)方策の中で漸近的に最小分散を達成する。
- 分散最小化MDPの本質的取り扱いにくさにもかかわらず、効率的設計の特徴付けは簡潔な凸最適化問題に還元される。
- TAR条件の下で非パラメトリックMLEは一貫的であり、真の定常状態報酬差に収束することが保証される。
- MLEの漸近的分散は、マルティングール差の直交性と大数の法則の議論により、すべてのTAR方策の中で最小の値に収束することが示された。
- オンライン設計は確率的に真の処置効果に収束し、状態行動ペairの経験的頻度の収束速度は$O(1/n)$で上限が与えられる。
- モデルの誤指定に対してもロバストであり、パラメトリックな仮定がなくても弱い正則性条件(TAR)の下で一貫性と効率性が保たれる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。