Skip to main content
QUICK REVIEW

[Paper Review] Robust Design and Evaluation of Predictive Algorithms under Unobserved Confounding

Ashesh Rambachan, Amanda Coston|arXiv (Cornell University)|Dec 19, 2022
Financial Distress and Bankruptcy Prediction4 citations
TL;DR

This paper proposes a robust framework for designing and evaluating predictive algorithms under unobserved confounding in selectively observed data, using flexible, covariate-dependent bounds on potential outcome differences. It develops debiased machine learning estimators for performance bounds—such as mean squared error, accuracy, and false positive rates—enabling sensitivity analysis that accounts for unmeasured confounders without strong parametric assumptions.

ABSTRACT

Predictive algorithms inform consequential decisions in settings with selective labels: outcomes are observed only for units selected by past decision makers. This creates an identification problem under unobserved confounding -- when selected and unselected units differ in unobserved ways that affect outcomes. We propose a framework for robust design and evaluation of predictive algorithms that bounds how much outcomes may differ between selected and unselected units with the same observed characteristics. These bounds formalize common empirical strategies including proxy outcomes and instrumental variables. Our estimators work across bounding strategies and performance measures such as conditional likelihoods, mean square error, and true/false positive rates. Using administrative data from a large Australian financial institution, we show that varying confounding assumptions substantially affects credit risk predictions and fairness evaluations across income groups.

Motivation & Objective

  • Address the challenge of evaluating predictive algorithms when outcomes are selectively observed due to human decision-making, with unobserved confounders affecting both selection and outcomes.
  • Overcome limitations of existing methods that assume unconfounded selection or rely on ad hoc imputation strategies for missing data.
  • Develop a general identification framework that formalizes common empirical strategies—like proxy outcomes and instrumental variables—within a partial identification approach.
  • Enable robust evaluation of predictive performance metrics (e.g., MSE, accuracy, false positive rates) under uncertainty about unobserved confounding.
  • Provide a sensitivity analysis framework that allows researchers to assess how conclusions vary across plausible assumptions about unobserved confounding.

Proposed method

  • Formalize identification assumptions by bounding the average potential outcome difference between selected and unselected units, conditionally on observed covariates and identified nuisance parameters.
  • Use a general class of restrictions on the mean difference in potential outcomes, extending prior sensitivity analysis methods to be conditional on covariates.
  • Develop debiased machine learning estimators for bounds on predictive performance estimands, ensuring asymptotic normality and valid inference.
  • Incorporate flexible nonparametric estimation of nuisance functions (e.g., propensity scores, outcome regressions) using machine learning methods.
  • Apply double/debiased estimation to reduce bias in the bounds estimation, enabling valid confidence intervals under weak regularity conditions.
  • Construct bounds on performance metrics such as mean squared error, accuracy, and false positive rates using the identified potential outcome differences.

Experimental results

Research questions

  • RQ1How can predictive algorithms be robustly evaluated when outcomes are only observed for selected units due to unobserved confounding?
  • RQ2What are the implications of different assumptions about unobserved confounding on default risk predictions and fairness metrics in credit scoring?
  • RQ3Can a unified framework be developed to bound performance metrics like MSE and accuracy under partial identification from selectively observed data?
  • RQ4How do covariate-dependent bounds on unobserved confounding improve the precision and interpretability of performance evaluations compared to unconditional bounds?
  • RQ5To what extent do standard imputation strategies (e.g., proxy outcomes, instrumental variables) fit within a formal partial identification framework?

Key findings

  • The proposed method produces informative bounds on predictive performance metrics by flexibly incorporating observed covariates in the identification assumptions.
  • Debiased machine learning estimators for the bounds achieve good finite-sample performance, with empirical coverage rates of 93.4–95.0% for 95% confidence intervals across sample sizes of 500 to 2500.
  • In an Australian financial dataset, varying assumptions about unobserved confounding led to meaningful shifts in default risk predictions and disparities in credit score evaluations across sensitive groups.
  • The framework allows for sensitivity analysis that systematically explores how conclusions change under different plausible levels of unobserved confounding.
  • The method generalizes and formalizes common empirical strategies such as proxy outcomes and instrumental variables within a single, coherent partial identification framework.
  • The approach enables robust evaluation of predictive algorithms even when selection is not conditionally unconfounded, offering a practical alternative to strong ignorability assumptions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.