[Paper Review] Estimating and accounting for unobserved covariates in high dimensional correlated data
This paper proposes CBCV and CorrConf, computationally efficient methods to estimate the number and structure of latent confounding factors in high-dimensional correlated data—such as longitudinal or multi-tissue 'omics' data—where standard methods fail due to non-i.i.d. residuals. The methods achieve provably accurate estimation and correction of latent factors, improving downstream inference on known covariates.
Many high dimensional and high-throughput biological datasets have complex sample correlation structures, which include longitudinal and multiple tissue data, as well as data with multiple treatment conditions or related individuals. These data, as well as nearly all high-throughput `omic' data, are influenced by technical and biological factors unknown to the researcher, which, if unaccounted for, can severely obfuscate estimation and inference on effects due to the known covariate of interest. We therefore developed CBCV and CorrConf: provably accurate and computationally efficient methods to choose the number of and estimate latent confounding factors present in high dimensional data with correlated or nonexchangeable residuals. We demonstrate each method's superior performance compared to other state of the art methods by analyzing simulated multi-tissue gene expression data and identifying sex-associated DNA methylation sites in a real, longitudinal twin study. As far as we are aware, these are the first methods to estimate the number of and correct for latent confounding factors in data with correlated or nonexchangeable residuals. An R-package is available for download at https://github.com/chrismckennan/CorrConf.
Motivation & Objective
- To address the challenge of unobserved confounders in high-dimensional biological data with complex correlation structures, such as longitudinal, multi-tissue, or related-individual studies.
- To develop a method that accurately estimates both the number and structure of latent confounding factors when residuals are correlated or nonexchangeable, a scenario where existing methods fail.
- To enable accurate downstream inference on known covariates of interest by correcting for latent factors without assuming independent and identically distributed residuals.
- To provide a provably accurate and computationally efficient solution for data with heterogeneous residual variances and structured correlation, such as those arising from repeated measures or batch effects.
Proposed method
- Proposes CBCV (Cross-Validation-based Criterion) and CorrConf (Correlation-Confounding) as novel methods to jointly estimate the number of latent factors and their structure in high-dimensional data with correlated residuals.
- Uses a likelihood-based framework that accounts for non-i.i.d. residual structures through a structured variance-covariance matrix, including group-specific and time-specific components.
- Incorporates a two-stage estimation: first estimating the number of latent factors via a penalized likelihood criterion sensitive to correlation structure, then estimating the factor loadings using a corrected generalized least squares approach.
- Employs a projection-based estimator to decouple the effects of observed and unobserved covariates, ensuring consistent estimation of the coefficient of interest even when latent factors are unobserved.
- Derives asymptotic normality of the estimated coefficients under a double asymptotic regime where both sample size and dimension grow, establishing theoretical validity.
- Uses a rotation-invariant estimator for the latent factor space and applies a shrinkage correction to the residual variance estimator to avoid overfitting in high-dimensional settings.
Experimental results
Research questions
- RQ1How can we accurately estimate the number of latent confounding factors in high-dimensional data with correlated or nonexchangeable residuals?
- RQ2What is the impact of ignoring residual correlation on standard latent factor estimation methods, and how does this affect downstream inference on known covariates?
- RQ3Can we develop a method that jointly estimates the number and structure of latent factors in data with complex correlation structures, such as longitudinal or multi-tissue data?
- RQ4To what extent does correcting for latent factors improve the power and accuracy of detecting true biological effects in high-throughput 'omics' data?
Key findings
- CBCV and CorrConf outperform existing state-of-the-art methods in estimating the number of latent factors and correcting for confounding in simulated multi-tissue gene expression data.
- In a real longitudinal twin study with 8×10⁵ CpG sites, the methods successfully identified sex-associated DNA methylation sites with higher accuracy than competing approaches.
- The methods are the first to provide provably accurate estimation of latent factors in data with correlated or nonexchangeable residuals, a critical gap in current statistical methodology.
- Theoretical analysis shows that the estimator for the coefficient of interest is asymptotically normal with a variance-covariance structure that accounts for unobserved confounding, ensuring valid inference.
- The methods avoid the common pitfall of overestimating the number of latent factors in correlated data, which plagues standard methods like parallel analysis and bi-cross validation.
- An R package named CorrConf is publicly available at https://github.com/chrismckennan/CorrConf, enabling broad reproducibility and application in real-world studies.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.