[论文解读] Inference with Transposable Data: Modeling the Effects of Row and Column Correlations
该论文提出了一种方法,通过使用可转置正则化协方差估计来建模并校正行和列之间的相关性,从而提升大规模可转置矩阵数据的推理性能。通过利用估计的行和列协方差矩阵对数据进行球化变换,该方法恢复了正确的零分布和独立性,显著提高了微阵列研究中的统计功效,并改善了错误发现率的估计。
We consider the problem of large-scale inference on the row or column variables of data in the form of a matrix. Often this data is transposable, meaning that both the row variables and column variables are of potential interest. An example of this scenario is detecting significant genes in microarrays when the samples or arrays may be dependent due to underlying relationships. We study the effect of both row and column correlations on commonly used test-statistics, null distributions, and multiple testing procedures, by explicitly modeling the covariances with the matrix-variate normal distribution. Using this model, we give both theoretical and simulation results revealing the problems associated with using standard statistical methodology on transposable data. We solve these problems by estimating the row and column covariances simultaneously, with transposable regularized covariance models, and de-correlating or sphering the data as a pre-processing step. Under reasonable assumptions, our method gives test statistics that follow the scaled theoretical null distribution and are approximately independent. Simulations based on various models with structured and observed covariances from real microarray data reveal that our method offers substantial improvements in two areas: 1) increased statistical power and 2) correct estimation of false discovery rates.
研究动机与目标
- 为解决标准统计方法在具有行和列相关性的可转置矩阵数据上应用时的局限性。
- 在大规模推断中,特别是微阵列研究中,对行和列相关性进行建模并加以校正。
- 在依赖结构下提高检验统计量和零分布的准确性。
- 在高维、相关数据中提升统计功效,并实现错误发现率的正确估计。
- 开发一种实用的预处理方法——通过可转置正则化协方差估计进行球化变换,以实现稳健的推断。
提出的方法
- 使用均值受限的矩阵变正态分布对数据进行建模,以显式捕捉行和列之间的相关性。
- 利用可转置正则化协方差模型,同时估计行和列协方差矩阵。
- 通过预乘以估计行协方差矩阵的逆平方根和后乘以估计列协方差矩阵的逆平方根,对数据应用球化变换以去除相关性。
- 利用变换后的数据计算检验统计量,使其服从理论上的零分布且近似独立。
- 利用矩阵变正态分布的特征函数,推导在零假设下检验统计量的渐近分布。
- 通过模拟和真实微阵列数据验证该方法,与标准方法进行性能比较。
实验结果
研究问题
- RQ1行和列相关性如何影响可转置数据中标准检验统计量的零分布和统计功效?
- RQ2忽略行和列相关性对多重检验程序和错误发现率估计有何影响?
- RQ3同时估计行和列协方差是否能提高检验统计量和零分布的有效性?
- RQ4通过估计协方差对数据进行球化在多大程度上能恢复独立性并正确缩放检验统计量?
- RQ5在真实微阵列数据上,该方法与标准方法相比,在统计功效和FDR估计方面表现如何?
主要发现
- 所提出的方法即使在复杂的相关结构下,也能使检验统计量恢复到服从理论上的零分布。
- 经球化变换后的数据所导出的检验统计量近似独立,从而支持有效的多重检验校正。
- 模拟结果表明,与忽略相关性的标准方法相比,该方法在统计功效上实现了显著提升。
- 在相关性较强时,该方法在错误发现率估计方面显著更准确。
- 在真实微阵列数据(包括Cardio和Leukemia数据集)上,该方法优于现有方法,通过校正t统计量的过度离散性实现了改进。
- 理论分析证实,在零假设下,变换后的检验统计量服从缩放t分布,其缩放因子取决于估计的协方差结构。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。