Skip to main content
QUICK REVIEW

[论文解读] Off-policy Policy Evaluation For Sequential Decisions Under Unobserved Confounding

Hongseok Namkoong, Ramtin Keramati|arXiv (Cornell University)|Mar 12, 2020
Health Systems, Economic Evaluations, Quality of Life参考文献 55被引用 18
一句话总结

该论文提出了一种在未观测混杂因素下对序列决策进行稳健的离策略评估(OPE)的方法,重点关注仅有一次决策受潜在混杂因素影响的情境。通过函数凸对偶性和损失最小化过程,该方法计算出统计一致的最坏情况性能边界,即使在标准OPE方法因未观测因素而产生偏差时,也能实现可靠的策略评估。

ABSTRACT

When observed decisions depend only on observed features, off-policy policy evaluation (OPE) methods for sequential decision making problems can estimate the performance of evaluation policies before deploying them. This assumption is frequently violated due to unobserved confounders, unrecorded variables that impact both the decisions and their outcomes. We assess robustness of OPE methods under unobserved confounding by developing worst-case bounds on the performance of an evaluation policy. When unobserved confounders can affect every decision in an episode, we demonstrate that even small amounts of per-decision confounding can heavily bias OPE methods. Fortunately, in a number of important settings found in healthcare, policy-making, operations, and technology, unobserved confounders may primarily affect only one of the many decisions made. Under this less pessimistic model of one-decision confounding, we propose an efficient loss-minimization-based procedure for computing worst-case bounds, and prove its statistical consistency. On two simulated healthcare examples---management of sepsis patients and developmental interventions for autistic children---where this is a reasonable model of confounding, we demonstrate that our method invalidates non-robust results and provides meaningful certificates of robustness, allowing reliable selection of policies even under unobserved confounding.

研究动机与目标

  • 为解决在序列决策的离策略评估(OPE)中,由于隐藏变量影响决策和结果而导致标准OPE方法失效的严峻挑战,提出应对未观测混杂因素的机制。
  • 开发一种敏感性分析框架,量化未观测混杂因素对OPE估计的影响,特别是在仅一个决策受混杂影响的场景中。
  • 在受限未观测混杂因素下,提供评估策略性能的统计一致且可计算的最坏情况边界,确保即使存在混杂因素,也能实现可靠的策略选择。
  • 在真实世界医疗应用(如败血症管理与自闭症干预)中展示该方法的实际效用,这些场景中混杂因素虽可能存在但仅限于一个决策点。
  • 通过利用凸对偶性和重要性采样,将现有单决策混杂模型扩展至序列设置,确保计算可行性与理论一致性。

提出的方法

  • 提出一种受限未观测混杂模型,其中序列中仅一个决策受潜在变量影响,相较于全序列混杂的假设更为宽松。
  • 利用函数凸对偶性,推导最坏情况性能估计问题的对偶松弛,将其转化为凸优化问题。
  • 提出一种基于损失最小化的程序,通过使用神经网络和反向传播求解加权平方损失最小化问题,以计算最坏情况边界。
  • 采用重要性采样校正观测特征,确保边界与观测行为策略一致,保持与观察数据的一致性。
  • 使用四层ReLU神经网络配合Adam优化器来估计损失最小化目标,无观测混杂因素时则使用逻辑回归建模边际行为策略。
  • 证明了最坏情况边界估计器的经验近似具有统计一致性,确保随着样本量增加而收敛。

实验结果

研究问题

  • RQ1未观测混杂因素——特别是当仅影响序列中一个决策时——如何影响标准离策略评估(OPE)方法的可靠性?
  • RQ2我们能否在序列决策中,针对评估策略计算出在受限未观测混杂因素下依然有效的最坏情况性能边界?
  • RQ3在真实医疗应用中,我们的方法在多大程度上能验证OPE结果对未观测混杂因素的稳健性?
  • RQ4与朴素敏感性分析相比,该方法在设计敏感性和检测混杂条件下策略优越性方面表现如何?
  • RQ5在存在未观测混杂因素的情况下,所提出的损失最小化程序在计算最坏情况边界时的统计一致性如何?

主要发现

  • 在败血症管理模拟中,标准OPE在轻微混杂(Γ=2.0)下将自适应策略的性能高估了高达20%,表明出现了虚假的策略优越性。
  • 所提方法检测到即使在Γ=2.0时,OPE估计值也可能因未观测混杂因素而产生不可接受的偏差,从而在存在未观测混杂因素时使标准OPE的结论无效。
  • 在第二个场景中,自适应策略确实更优,该方法成功认证其优势可达Γ=4.2,表明对非平凡水平的混杂因素具有鲁棒性。
  • 所提方法的设计敏感度达到Γ=2.78,显著优于朴素方法(Γ=1.32),表明其在混杂因素存在时更强的策略性能验证能力。
  • 该方法成功识别出,由于未观测混杂因素,OPE可能错误地显示自适应策略优于非自适应策略,即使真实性能相反。
  • 实证结果表明,损失最小化程序在中等时域、高维连续状态和奖励空间中具有统计一致性与有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。