[Paper Review] Causal Mediation Analysis Leveraging Multiple Types of Summary Statistics Data
This paper proposes CaMMEL, a Bayesian multivariate mediation framework that leverages GWAS and eQTL summary statistics to identify causal genes mediating SNP-disease relationships without individual-level data. By modeling high-dimensional, correlated genetic variants and projecting out confounding effects via LD blocks, CaMMEL achieves superior accuracy in identifying true mediators even under polygenic bias and missing data, outperforming existing methods in simulations and real-world scenarios.
Summary statistics of genome-wide association studies (GWAS) teach causal relationship between millions of genetic markers and tens and thousands of phenotypes. However, underlying biological mechanisms are yet to be elucidated. We can achieve necessary interpretation of GWAS in a causal mediation framework, looking to establish a sparse set of mediators between genetic and downstream variables, but there are several challenges. Unlike existing methods rely on strong and unrealistic assumptions, we tackle practical challenges within a principled summary-based causal inference framework. We analyzed the proposed methods in extensive simulations generated from real-world genetic data. We demonstrated only our approach can accurately redeem causal genes, even without knowing actual individual-level data, despite the presence of competing non-causal trails.
Motivation & Objective
- To address the lack of interpretability in GWAS by identifying causal genes that mediate SNP effects on complex diseases.
- To overcome limitations of existing methods that rely on strong assumptions and fail under polygenic bias, missing mediators, and confounding.
- To develop a principled, summary-based causal inference framework that integrates multiple types of genetic association statistics (GWAS and eQTL).
- To enable accurate identification of causal mediators in high-dimensional, collinear genetic data without access to individual-level genotypes.
- To provide a robust, scalable method for causal mediation analysis applicable to biobank-scale genetic studies.
Proposed method
- Proposes a Bayesian multivariate model to jointly analyze GWAS and eQTL summary statistics, modeling SNP effects on both outcome and mediator (gene) variables.
- Uses sparse Bayesian variable selection to handle high-dimensional collinearity from LD, enabling selection of true causal mediators.
- Introduces two key operational steps: factorization (CaMMEL-fact) and projection (CaMMEL-proj) to disentangle genetic effects from non-genetic confounding.
- Applies LD block decomposition (1,703 independent blocks) to project correlation structures and isolate spurious confounding effects.
- Employs a hierarchical prior structure to model unmediated (direct) effects and mediators simultaneously, improving identifiability.
- Uses a likelihood-based inference framework that accounts for uncertainty in summary statistics and enables posterior inference on causal effects.
Experimental results
Research questions
- RQ1Can a summary-based causal mediation framework accurately identify causal genes mediating SNP effects on complex diseases without individual-level data?
- RQ2How does the method perform under polygenic bias, missing mediators, and unmeasured confounding?
- RQ3Can the framework distinguish between direct genetic effects and spurious confounding due to non-genetic factors?
- RQ4How does the inclusion of multiple mediators in a Bayesian framework improve identification over univariate gene-by-gene methods?
- RQ5To what extent do LD block projections and factorization steps enhance robustness to confounding and collinearity?
Key findings
- CaMMEL-proj and CaMMEL-fact significantly outperform other methods in identifying true causal mediators when confounding or polygenic bias is present.
- Under 50% missing genes and directional pleiotropy, CaMMEL methods maintain high AUPRC, while univariate methods suffer severe accuracy loss.
- The naive inference algorithm on the CaMMEL model performs poorly due to overfitting to unmediated effects, highlighting the need for structured regularization.
- Even without explicit modeling of confounders, CaMMEL-proj successfully separates non-genetic confounding from genetic effects by projecting onto independent LD blocks.
- In simulations with full gene observation, CaMMEL-proj achieves near-optimal performance when no confounding exists, confirming robustness.
- The method demonstrates superior power and precision in identifying causal genes even when only a single gene is truly causal among 100 candidates.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.