[Paper Review] Policy Learning under Biased Sample Selection
This paper proposes robust policy learning under biased sample selection by modeling sampling bias via a $Γ$-biased sampling framework, where unobserved heterogeneity is bounded. It derives optimal max-min and max-min gain policies that depend on both the conditional average treatment effect (CATE) and conditional value-at-risk (CVaR), enabling performance guarantees under worst-case external validity failures without plug-in estimation of nuisance parameters.
Practitioners often use data from a randomized controlled trial to learn a treatment assignment policy that can be deployed on a target population. A recurring concern in doing so is that, even if the randomized trial was well-executed (i.e., internal validity holds), the study participants may not represent a random sample of the target population (i.e., external validity fails)--and this may lead to policies that perform suboptimally on the target population. We consider a model where observable attributes can impact sample selection probabilities arbitrarily but the effect of unobservable attributes is bounded by a constant, and we aim to learn policies with the best possible performance guarantees that hold under any sampling bias of this type. In particular, we derive the partial identification result for the worst-case welfare in the presence of sampling bias and show that the optimal max-min, max-min gain, and minimax regret policies depend on both the conditional average treatment effect (CATE) and the conditional value-at-risk (CVaR) of potential outcomes given covariates. To avoid finite-sample inefficiencies of plug-in estimates, we further provide an end-to-end procedure for learning the optimal max-min and max-min gain policies that does not require the separate estimation of nuisance parameters.
Motivation & Objective
- To address the challenge of external validity failure in policy learning from randomized controlled trials (RCTs) due to non-representative study populations.
- To develop policy learning methods that remain robust when study participants are not representative of the target population due to sampling bias.
- To provide performance guarantees under worst-case sampling bias, where unobserved factors influence selection with bounded impact.
- To avoid finite-sample inefficiencies from plug-in estimation of nuisance parameters in policy learning.
- To derive and implement end-to-end procedures for learning optimal max-min and max-min gain policies under partial identification.
Proposed method
- Models sampling bias using the $\Gamma$-biased sampling framework, allowing arbitrary covariate-dependent selection but bounded unobserved heterogeneity.
- Derives partial identification bounds for worst-case welfare under sampling bias, showing that optimal policies depend on both CATE and CVaR of potential outcomes.
- Proposes an end-to-end learning procedure that bypasses separate estimation of nuisance parameters by directly optimizing over the worst-case distribution.
- Uses a minimax formulation to learn policies that maximize worst-case expected outcome across all plausible target distributions consistent with observed data and $\Gamma$-bias.
- Employs a quantile-based construction of the worst-case conditional outcome distribution, with weights $\Gamma^{-1}$ and $\Gamma$ assigned to outcomes above and below a $\zeta(\Gamma) = \frac{1}{\Gamma+1}$-quantile.
- Establishes that the optimal policy is determined by a combination of CATE and CVaR, ensuring robustness to unobserved selection effects.
Experimental results
Research questions
- RQ1How can we learn treatment assignment policies that maintain performance guarantees when the study population is not representative of the target population due to sampling bias?
- RQ2What is the worst-case welfare achievable under $\Gamma$-biased sampling, and how can it be characterized using observable data?
- RQ3How do the optimal max-min and max-min gain policies depend on the conditional average treatment effect (CATE) and conditional value-at-risk (CVaR)?
- RQ4Can we avoid finite-sample inefficiencies from plug-in estimation of nuisance parameters in robust policy learning?
- RQ5What is the structure of the worst-case distribution over potential outcomes that maximizes the lower bound on policy performance?
Key findings
- The optimal max-min policy depends on both the conditional average treatment effect (CATE) and the conditional value-at-risk (CVaR) of potential outcomes given covariates.
- The optimal max-min gain policy also depends on CATE and CVaR, reflecting a trade-off between average performance and worst-case robustness.
- The worst-case distribution of potential outcomes under $\Gamma$-biased sampling assigns weights $\Gamma^{-1}$ and $\Gamma$ to outcomes above and below the $\frac{1}{\Gamma+1}$-quantile, respectively.
- The proposed end-to-end procedure avoids separate estimation of nuisance parameters, improving finite-sample efficiency compared to plug-in methods.
- Partial identification bounds for worst-case welfare are derived, showing that performance guarantees hold across all distributions consistent with $\Gamma$-biased sampling.
- The method ensures that learned policies maintain performance under external validity failures due to unobserved selection bias, even when the target population differs from the study population.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.