[Paper Review] Factor analysis in high dimensional biological data with dependent observations
This paper proposes a novel high-dimensional factor analysis framework for biological data with dependent observations and heterogeneous signal strengths across factors. It introduces a robust estimator for the number of factors that overcomes eigenvalue shadowing and pervasive factor assumption biases, enabling accurate subspace recovery, denoising, and latent factor interpretation in longitudinal, multi-tissue, and multi-treatment data.
Factor analysis is a critical component of high dimensional biological data analysis. However, modern biological data contain two key features that irrevocably corrupt existing methods. First, these data, which include longitudinal, multi-treatment and multi-tissue data, contain samples that break critical independence requirements necessary for the utilization of prevailing methods. Second, biological data contain factors with large, moderate and small signal strengths, and therefore violate the ubiquitous "pervasive factor" assumption essential to the performance of many methods. In this work, I develop a novel statistical framework to perform factor analysis and interpret its results in data with dependent observations and factors whose signal strengths span several orders of magnitude. I then prove that my methodology can be used to solve many important and previously unsolved problems that routinely arise when analyzing dependent biological data, including high dimensional covariance estimation, subspace recovery, latent factor interpretation and data denoising. Additionally, I show that my estimator for the number of factors overcomes both the notorious "eigenvalue shadowing" problem, as well as the biases due to the pervasive factor assumption that plague existing estimators. Simulated and real data demonstrate the superior performance of my methodology in practice.
Motivation & Objective
- Address the critical gap in existing factor analysis methods that assume independent samples and pervasive factors, both of which are violated in modern high-dimensional biological data.
- Develop a statistically sound framework for factor analysis under dependent observations, where the error matrix rows may be correlated but have bounded eigenvalues.
- Enable accurate estimation of the number of latent factors, loadings, and factors even when signal strengths span multiple orders of magnitude.
- Overcome the limitations of existing estimators that fail to recover moderate and weak factors due to the pervasive factor assumption and eigenvalue shadowing.
- Facilitate downstream biological inference by enabling data denoising, covariance estimation, and interpretation of latent biological variation in complex data structures.
Proposed method
- Propose a general factor model where the observed data matrix $\bm{Y}$ is decomposed into $\bm{L}\bm{C}^\top + \bm{E}$, allowing for dependent error structure with bounded eigenvalues in $\bm{V}_g$.
- Introduce a two-step estimation procedure: first, estimate the factor space using a modified principal component analysis on a transformed data matrix to account for dependence.
- Use a rotation-based estimator for the number of factors $K$ that leverages eigenvalue ratios and is robust to signal strength heterogeneity.
- Apply a unitary transformation and leverage score control via random matrix theory to ensure consistent estimation of factor loadings and scores under dependence.
- Incorporate a diagonal correction for the covariance matrix $\hat{\bm{\Sigma}}$ to stabilize variance estimation in high-dimensional settings.
- Use a co-factor expansion argument and sub-exponential tail bounds to derive asymptotic normality and consistency of the estimated factors and loadings.
Experimental results
Research questions
- RQ1How can factor analysis be reliably performed in high-dimensional biological data when samples are dependent, violating the i.i.d. assumption of standard methods?
- RQ2Can a factor model be developed that accurately estimates the number of factors when signal strengths vary widely (large, moderate, and small), avoiding the pitfalls of the pervasive factor assumption?
- RQ3To what extent does the proposed estimator for $K$ overcome the eigenvalue shadowing problem common in high-dimensional data?
- RQ4Can the proposed method consistently recover latent factors and loadings under dependence and heterogeneous signal strengths, and how does it compare to existing PCA-based estimators?
- RQ5Does the framework enable effective data denoising and improved inference in eQTL/meQTL studies despite complex dependence structures?
Key findings
- The proposed estimator for the number of factors $K$ is consistent and robust to both eigenvalue shadowing and the pervasive factor assumption, outperforming existing methods in simulations and real data.
- The method successfully recovers moderate and weak factors—previously missed by standard estimators—by avoiding reliance on the $\lambda_K \to \infty$ assumption.
- The estimator for the factor space achieves $O_P(n^{-1/2})$ convergence rate, even under dependence, due to careful control of leverage scores and eigenvalue ratios.
- The framework enables accurate high-dimensional covariance estimation and subspace recovery, with theoretical guarantees under a general dependence structure.
- Data denoising using the estimated factors leads to improved downstream inference in eQTL and meQTL studies, as demonstrated on real multi-tissue and longitudinal datasets.
- Theoretical analysis shows that the rotation-based estimator for $K$ is consistent even when $\lambda_K \lesssim 1$, a regime where existing methods fail.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.