[Paper Review] Accounting for hidden common causes when inferring cause and effect from observational data
This paper proposes using L2-regularized multiple linear regression with a realized relationship matrix (RRM) to account for hidden common causes—such as family relatedness—in genome-wide association studies (GWAS). By modeling genetic similarity through the RRM, the method effectively infers hidden confounders, calibrates p-values, and enables accurate causal inference even when unmeasured population structure biases associations.
Identifying causal relationships from observation data is difficult, in large part, due to the presence of hidden common causes. In some cases, where just the right patterns of conditional independence and dependence lie in the data---for example, Y-structures---it is possible to identify cause and effect. In other cases, the analyst deliberately makes an uncertain assumption that hidden common causes are absent, and infers putative causal relationships to be tested in a randomized trial. Here, we consider a third approach, where there are sufficient clues in the data such that hidden common causes can be inferred.
Motivation & Objective
- To address the challenge of hidden common causes—such as population structure or family relatedness—in observational data when inferring causal relationships.
- To develop a method that enables accurate causal inference in GWAS despite the presence of unmeasured confounders that induce spurious associations between SNPs and traits.
- To demonstrate that p-values remain calibrated when using L2-regularized regression with a realized relationship matrix, even when hidden confounders are present.
- To show that the method effectively infers hidden common causes by leveraging the correlation structure among SNPs, thereby blocking d-connecting paths in the causal graph.
Proposed method
- Uses L2-regularized multiple linear regression to model the trait as a function of a test SNP and all other SNPs (termed 'similarity SNPs').
- Models the joint distribution of the trait as a linear mixed model: $\mathbf{y} \sim \mathcal{N}(\mathbf{1}\mu + \mathbf{x}^*\beta^*; \sigma_e^2\mathbf{I} + \sigma_g^2\mathbf{XX}^T)$, where $\mathbf{XX}^T$ is the realized relationship matrix (RRM).
- Imposes a hierarchical prior on the regression coefficients of similarity SNPs, assuming $\beta_i \sim \mathcal{N}(0, \sigma_g^2)$, which induces shrinkage and enables estimation of the variance component $\sigma_g^2/\sigma_e^2$.
- Estimates model parameters via restricted maximum likelihood (REML), with the ratio $\sigma_g^2/\sigma_e^2$ determined by grid search.
- Computes p-values using an F-test for $\beta^* = 0$, with the RRM capturing the influence of hidden confounders through SNP similarity.
- Uses a single, shared estimate of $\sigma_g^2/\sigma_e^2$ across all SNPs to improve computational efficiency without sacrificing accuracy.
Experimental results
Research questions
- RQ1Can p-values remain calibrated in GWAS when hidden common causes such as family relatedness are present?
- RQ2To what extent can the realized relationship matrix (RRM) derived from all SNPs effectively infer and block the influence of unmeasured hidden confounders?
- RQ3How does the performance of L2-regularized regression compare to univariate regression in the presence of hidden confounders?
- RQ4Does the method remain robust when there is a direct causal effect from the hidden common cause to the trait (e.g., environmental differences across populations)?
Key findings
- In the absence of a direct path from the hidden variable to the trait, p-values from L2-regularized regression were well-calibrated across a wide range of family relatedness levels, causal SNP counts, and effect strengths.
- When a direct path from the hidden variable to the trait was introduced (e.g., due to population-level environmental effects), p-values remained calibrated due to the effective inference of the hidden confounder via SNP similarity.
- The realized relationship matrix (RRM) constructed from all SNPs closely approximated the true underlying confounding structure, even when only causal SNPs were used to generate the data.
- The method achieved calibration across 450 synthetic GWAS data sets with varying parameters, including 50% to 90% family relatedness, 10 to 1000 causal SNPs, and heritability values from 0.1 to 0.6.
- The use of a single, shared estimate of $\sigma_g^2/\sigma_e^2$ across all SNPs maintained p-value calibration with minimal computational cost.
- The results demonstrate that the method effectively infers hidden common causes by leveraging the correlation structure among SNPs, thereby blocking d-connecting paths and enabling valid causal inference.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.