[Paper Review] Generalizing causal inferences from randomized trials: counterfactual and graphical identification
This paper develops a counterfactual and graphical framework to identify conditions under which causal effects from randomized trials can be generalized to a target population of trial-eligible individuals. It introduces a hypothetical intervention to scale up trial engagement (invitation and participation) and treatment assignment, showing that generalizability requires modeling direct effects of engagement on outcomes and using g-formula or inverse probability weighting for identification when such effects exist or are absent.
When engagement with a randomized trial is driven by factors that affect the outcome or when trial engagement directly affects the outcome independent of treatment, the average treatment effect among trial participants is unlikely to generalize to a target population. In this paper, we use counterfactual and graphical causal models to examine under what conditions we can generalize causal inferences from a randomized trial to the target population of trial-eligible individuals. We offer an interpretation of generalizability analyses using the notion of a hypothetical intervention to "scale-up" trial engagement to the target population. We consider the interpretation of generalizability analyses when trial engagement does or does not directly affect the outcome, highlight connections with censoring in longitudinal studies, and discuss identification of the distribution of counterfactual outcomes via g-formula computation and inverse probability weighting. Last, we show how the methods can be extended to address time-varying treatments, non-adherence, and censoring.
Motivation & Objective
- To address the challenge of generalizing causal effects from randomized trials to a target population when trial participants are not representative of all eligible individuals.
- To formally model the distinction between invitation to participate and actual participation in a trial, recognizing their potential direct effects on outcomes.
- To develop a framework for identifying the distribution of counterfactual outcomes in the target population using structural causal models and do-calculus.
- To extend generalizability methods to settings with time-varying treatments, non-adherence, and censoring.
- To clarify that generalizability analyses identify the effect of jointly scaling up engagement and treatment, not just treatment alone, when engagement affects outcomes.
Proposed method
- Uses potential outcomes and counterfactuals to define the target causal estimand in the population of trial-eligible individuals.
- Employs directed acyclic graphs (DAGs) with selection nodes to represent the processes of invitation (R) and participation (S) in a trial.
- Applies do-calculus to identify the effect of a joint intervention setting R=1, S=1, and Z=z to emulate scaling up trial engagement and treatment assignment.
- Derives identification formulas using the g-formula to compute the distribution of counterfactual outcomes under the hypothetical intervention.
- Uses inverse probability weighting (IPW) as an alternative identification strategy when the g-formula is infeasible or when data are missing.
- Extends the framework to handle time-varying treatments, non-adherence, and censoring by modeling interventions that enforce adherence and eliminate censoring.
Experimental results
Research questions
- RQ1Under what conditions can causal effects estimated in a randomized trial be generalized to a target population of trial-eligible individuals?
- RQ2How does the presence of direct effects of trial engagement (invitation and participation) on outcomes affect the identifiability of generalizable causal effects?
- RQ3What is the correct causal estimand when generalizing from a trial to a target population, and how is it identified using counterfactual and graphical models?
- RQ4How can generalizability be assessed when trial participation does not directly affect outcomes, and how does this differ from settings with direct engagement effects?
- RQ5Can the proposed framework be extended to settings with time-varying treatments, non-adherence, and censoring?
Key findings
- Generalizability from a trial to a target population requires modeling both invitation (R) and participation (S) as distinct processes that may directly affect outcomes.
- When engagement has direct effects on outcomes, generalizability analyses identify the effect of jointly scaling up engagement and assigning treatment, not just treatment alone.
- The distribution of counterfactual outcomes under joint intervention on R, S, and Z can be identified using the g-formula when the conditional independence assumptions hold.
- Inverse probability weighting provides an alternative identification strategy, particularly useful when the g-formula is computationally or statistically challenging.
- The framework naturally extends to settings with time-varying treatments, non-adherence, and censoring by modeling interventions that enforce adherence and eliminate censoring.
- The paper clarifies that prior work assuming conditional independence of S given X may be misinterpreted when engagement has direct effects; such assumptions are valid only when engagement effects are negligible.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.