[论文解读] Policy Learning under Biased Sample Selection
本文通过在未观测异质性受约束的 $\Gamma$-偏差采样框架下建模采样偏差,提出了一种在偏差样本选择下具有鲁棒性的策略学习方法。该方法推导出依赖于条件平均处理效应(CATE)和条件风险价值(CVaR)的最优极大极小策略与极大极小增益策略,从而在无需对干扰参数进行插补估计的情况下,实现了在最坏情况外部有效性失效下的性能保证。
Practitioners often use data from a randomized controlled trial to learn a treatment assignment policy that can be deployed on a target population. A recurring concern in doing so is that, even if the randomized trial was well-executed (i.e., internal validity holds), the study participants may not represent a random sample of the target population (i.e., external validity fails)--and this may lead to policies that perform suboptimally on the target population. We consider a model where observable attributes can impact sample selection probabilities arbitrarily but the effect of unobservable attributes is bounded by a constant, and we aim to learn policies with the best possible performance guarantees that hold under any sampling bias of this type. In particular, we derive the partial identification result for the worst-case welfare in the presence of sampling bias and show that the optimal max-min, max-min gain, and minimax regret policies depend on both the conditional average treatment effect (CATE) and the conditional value-at-risk (CVaR) of potential outcomes given covariates. To avoid finite-sample inefficiencies of plug-in estimates, we further provide an end-to-end procedure for learning the optimal max-min and max-min gain policies that does not require the separate estimation of nuisance parameters.
研究动机与目标
- 为解决由于研究人群不具备代表性而导致的随机对照试验(RCTs)中外部有效性失败的挑战。
- 开发在研究参与者因采样偏差而无法代表目标人群时仍保持鲁棒性的策略学习方法。
- 在未观测因素对选择的影响被有界的情况下,提供对最坏情况采样偏差的性能保证。
- 避免在策略学习中因对干扰参数进行插补估计而导致的有限样本效率低下。
- 推导并实现部分识别下最优极大极小与极大极小增益策略的端到端程序。
提出的方法
- 使用 $\Gamma$-偏差采样框架建模采样偏差,允许任意协变量依赖的选择,但对未观测异质性施加有界约束。
- 推导在采样偏差下最坏情况福利的部分识别边界,表明最优策略依赖于潜在结果的CATE与CVaR。
- 提出一种端到端学习程序,通过直接在最坏情况分布上进行优化,避免对干扰参数进行单独估计。
- 采用极小化极大公式,学习在所有与观测数据和 $\Gamma$-偏差一致的合理目标分布中,最大化最坏情况期望结果的策略。
- 基于分位数构造最坏情况下的条件结果分布,对高于和低于 $\zeta(\Gamma) = \frac{1}{\Gamma+1}$-分位数的结果分别赋予权重 $\Gamma^{-1}$ 和 $\Gamma$。
- 证明最优策略由CATE与CVaR的组合决定,从而对未观测的选择效应保持鲁棒性。
实验结果
研究问题
- RQ1如何学习治疗分配策略,使其在研究人群因采样偏差而不具备代表性时,仍能维持性能保证?
- RQ2在 $\Gamma$-偏差采样下,最坏情况下的福利可达水平是什么?如何利用可观测数据对其进行表征?
- RQ3最优极大极小策略与极大极小增益策略如何依赖于条件平均处理效应(CATE)与条件风险价值(CVaR)?
- RQ4我们能否避免在鲁棒策略学习中因对干扰参数进行插补估计而导致的有限样本效率低下?
- RQ5在潜在结果上最大化策略性能下界时,最坏情况分布的结构是什么?
主要发现
- 最优极大极小策略依赖于给定协变量的潜在结果的条件平均处理效应(CATE)与条件风险价值(CVaR)。
- 最优极大极小增益策略同样依赖于CATE与CVaR,体现了平均表现与最坏情况鲁棒性之间的权衡。
- 在 $\Gamma$-偏差采样下,潜在结果的最坏情况分布对高于 $\frac{1}{\Gamma+1}$-分位数的结果赋予权重 $\Gamma^{-1}$,对低于该分位数的结果赋予权重 $\Gamma$。
- 所提出的端到端程序避免了对干扰参数的单独估计,相比插补方法提高了有限样本效率。
- 推导出最坏情况福利的部分识别边界,表明性能保证在所有与 $\Gamma$-偏差采样一致的分布中均成立。
- 该方法确保所学习的策略在因未观测选择偏差导致的外部有效性失败下仍能保持性能,即使目标人群与研究人群不同。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。