[Paper Review] Inference with Transposable Data: Modeling the Effects of Row and Column Correlations
This paper proposes a method to improve large-scale inference on transposable matrix data by modeling and correcting for both row and column correlations using transposable regularized covariance estimation. By sphering the data via estimated row and column covariance matrices, the approach restores proper null distributions and independence, significantly boosting statistical power and improving false discovery rate estimation in microarray studies.
We consider the problem of large-scale inference on the row or column variables of data in the form of a matrix. Often this data is transposable, meaning that both the row variables and column variables are of potential interest. An example of this scenario is detecting significant genes in microarrays when the samples or arrays may be dependent due to underlying relationships. We study the effect of both row and column correlations on commonly used test-statistics, null distributions, and multiple testing procedures, by explicitly modeling the covariances with the matrix-variate normal distribution. Using this model, we give both theoretical and simulation results revealing the problems associated with using standard statistical methodology on transposable data. We solve these problems by estimating the row and column covariances simultaneously, with transposable regularized covariance models, and de-correlating or sphering the data as a pre-processing step. Under reasonable assumptions, our method gives test statistics that follow the scaled theoretical null distribution and are approximately independent. Simulations based on various models with structured and observed covariances from real microarray data reveal that our method offers substantial improvements in two areas: 1) increased statistical power and 2) correct estimation of false discovery rates.
Motivation & Objective
- To address the limitations of standard statistical methods when applied to transposable matrix data with correlated rows and columns.
- To model and correct for both row and column correlations in large-scale inference, particularly in microarray studies.
- To improve the accuracy of test statistics and null distributions under dependence structures.
- To enhance statistical power and enable correct estimation of false discovery rates in high-dimensional, correlated data.
- To develop a practical pre-processing method—sphering via transposable regularized covariance estimation—for robust inference.
Proposed method
- Model the data using the mean-restricted matrix-variate normal distribution to explicitly capture row and column correlations.
- Estimate row and column covariance matrices simultaneously using transposable regularized covariance models.
- Apply a sphering transformation to de-correlate the data by pre-multiplying with the inverse square root of the estimated row covariance and post-multiplying with the inverse square root of the estimated column covariance.
- Use the transformed data to compute test statistics that follow the theoretical null distribution and are approximately independent.
- Leverage the characteristic function of the matrix-variate normal distribution to derive the asymptotic distribution of the test statistics under the null hypothesis.
- Validate the method through simulations and real microarray data, comparing performance to standard approaches.
Experimental results
Research questions
- RQ1How do row and column correlations affect the null distribution and statistical power of standard test statistics in transposable data?
- RQ2What is the impact of ignoring row and column correlations on multiple testing procedures and false discovery rate estimation?
- RQ3Can simultaneous estimation of row and column covariances improve the validity of test statistics and null distributions?
- RQ4To what extent does sphering the data via estimated covariances restore independence and proper scaling of test statistics?
- RQ5How does the proposed method compare to standard approaches in terms of statistical power and FDR estimation on real microarray data?
Key findings
- The proposed method restores test statistics to follow the theoretical null distribution, even under complex correlation structures.
- Test statistics derived from sphered data are approximately independent, enabling valid multiple testing corrections.
- Simulations show substantial increases in statistical power compared to standard methods that ignore correlations.
- False discovery rate estimation is significantly more accurate under the proposed method, especially when correlations are strong.
- The method outperforms existing approaches on real microarray data, including the Cardio and Leukemia datasets, by correcting for over-dispersion in t-statistics.
- Theoretical analysis confirms that the transformed test statistics follow a scaled t-distribution under the null, with the scaling factor depending on the estimated covariance structure.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.