[Paper Review] Observational-Interventional Priors for Dose-Response Learning
This paper proposes a hierarchical Gaussian process prior that combines observational and interventional data to improve nonparametric dose-response curve estimation. By using observational data to inform a prior and refining it with a nonparametric affine transform from sparse interventional data, the method accelerates learning and improves accuracy, reducing average normalized absolute error by 0.06–0.08 in synthetic experiments on premature infant therapy data.
Controlled interventions provide the most direct source of information for learning causal effects. In particular, a dose-response curve can be learned by varying the treatment level and observing the corresponding outcomes. However, interventions can be expensive and time-consuming. Observational data, where the treatment is not controlled by a known mechanism, is sometimes available. Under some strong assumptions, observational data allows for the estimation of dose-response curves. Estimating such curves nonparametrically is hard: sample sizes for controlled interventions may be small, while in the observational case a large number of measured confounders may need to be marginalized. In this paper, we introduce a hierarchical Gaussian process prior that constructs a distribution over the dose-response curve by learning from observational data, and reshapes the distribution with a nonparametric affine transform learned from controlled interventions. This function composition from different sources is shown to speed-up learning, which we demonstrate with a thorough sensitivity analysis and an application to modeling the effect of therapy on cognitive skills of premature infants.
Motivation & Objective
- To address the challenge of learning dose-response curves when interventional data is scarce and expensive to collect.
- To improve nonparametric estimation of causal effects by combining observational data with limited interventional data.
- To develop a principled, hierarchical prior that leverages both data sources while maintaining robustness to model misspecification.
- To demonstrate the method’s performance on real-world data from a premature infant therapy study, where treatment levels were not randomized.
Proposed method
- The method employs a hierarchical Gaussian process prior that learns a base distribution over dose-response curves from observational data.
- It applies a nonparametric affine transformation to reshape the prior using interventional data, enabling efficient refinement with few intervention samples.
- The transformation is learned via a Gaussian process with a squared exponential covariance function, ensuring smoothness and flexibility.
- The approach uses back-door adjustment to estimate the interventional mean response, $ f(x) = \mathbb{E}[Y \mid do(X=x)] $, by marginalizing over measured confounders $ \mathbf{Z} $.
- The model is trained using a joint likelihood that combines observational and interventional data under the assumption that confounders $ \mathbf{Z} $ satisfy the back-door criterion.
- The method is evaluated via sensitivity analysis and synthetic interventional data generation from a real study on premature infants.
Experimental results
Research questions
- RQ1Can observational data be effectively leveraged to improve the estimation of dose-response curves when interventional data is limited?
- RQ2How does combining observational and interventional data via a hierarchical prior compare to using either data source alone in terms of estimation accuracy?
- RQ3How robust is the method to model misspecification, particularly when confounders are not fully measured or the back-door criterion is violated?
- RQ4To what extent does the method reduce estimation error compared to standard Gaussian process regression with only interventional data?
- RQ5How does the method perform across different strata of covariates, such as maternal education levels, in real-world data?
Key findings
- The method reduced average normalized absolute error by 0.06, 0.07, and 0.08 for the high school, college, and combined maternal education strata, respectively, when using 10 interventional samples per treatment level.
- In 82%, 89%, and 91% of runs, the method outperformed pure interventional Gaussian process regression in the high school, college, and combined strata, respectively.
- The method demonstrated improved convergence speed and robustness to model misspecification through sensitivity analysis.
- The posterior mean curve from the proposed method showed better fit to synthetic interventional data than standard GP regression with only interventional data.
- The approach effectively captured non-linear dose-response relationships in synthetic data generated from real infant development data, including varying ranges across maternal education strata.
- The method maintained strong performance even when confounding was not fully accounted for, suggesting practical utility in settings with unmeasured feedback effects.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.