[论文解读] Two-Sample Testing in High-Dimensional Models
本文提出了一种通用的、基于数据分割的高维两样本检验方法,适用于回归和高斯图模型等模型。通过结合样本分割、$\ell_1$-惩罚估计以及在多次分割中对p值进行聚合,该方法利用加权卡方分布的渐近零分布,实现了渐近有效的推断,即使在$p \gg n$时也能保持可靠性。与置换检验相比,该方法在统计功效上表现更优,并避免了单次分割带来的'p值彩票'问题。
We propose novel methodology for testing equality of model parameters between two high-dimensional populations. The technique is very general and applicable to a wide range of models. The method is based on sample splitting: the data is split into two parts; on the first part we reduce the dimensionality of the model to a manageable size; on the second part we perform significance testing (p-value calculation) based on a restricted likelihood ratio statistic. Assuming that both populations arise from the same distribution, we show that the restricted likelihood ratio statistic is asymptotically distributed as a weighted sum of chi-squares with weights which can be efficiently estimated from the data. In high-dimensional problems, a single data split can result in a "p-value lottery". To ameliorate this effect, we iterate the splitting process and aggregate the resulting p-values. This multi-split approach provides improved p-values. We illustrate the use of our general approach in two-sample comparisons of high-dimensional regression models ("differential regression") and graphical models ("differential network"). In both cases we show results on simulated data as well as real data from recent, high-throughput cancer studies.
研究动机与目标
- 解决在$p \gg n$的高维两样本问题中缺乏可靠显著性检验的问题。
- 开发一种适用于多种模型(包括高维回归和高斯图模型)的通用框架。
- 通过迭代分割和p值聚合,克服单次数据分割带来的不稳定性(即'p值彩票'问题)。
- 提供一种理论基础坚实的推断方法,其渐近零分布为加权卡方分布的和,且权重可从数据中估计。
- 实现对高通量生物数据中差异关系的统计检验,例如癌症基因组学中的差异网络。
提出的方法
- 将两个总体的数据分别划分为两部分:一部分用于模型筛选,另一部分用于假设检验。
- 在第一部分上使用$\ell_1$-惩罚最大似然估计,以选择活跃参数集,确保稀疏性和筛选性质。
- 在第二部分上,计算一个受限似然比统计量,用于比较非嵌套模型(个体参数空间 vs. 汇总参数空间)。
- 建立受限似然比统计量在零假设下的渐近分布为独立$\chi^2$变量的加权和,权重从数据中估计。
- 多次重复分割过程,并对p值进行聚合,以降低变异性和提高可靠性。
- 通过惩罚似然估计,将该方法应用于差异回归(高维线性模型)和差异网络(高斯图模型)分析。
实验结果
研究问题
- RQ1能否为高维模型($p \gg n$)开发一种通用的两样本检验框架?
- RQ2当经典似然比检验在高维情形下失效时,如何实现可靠的显著性检验?
- RQ3通过数据分割和p值聚合,能否缓解高维设定下单次分割p值的不稳定性?
- RQ4在非嵌套的高维模型中,受限似然比统计量的渐近零分布是什么?
- RQ5该方法能否检测出在基因调控网络或癌症亚型间回归效应中的生物学上具有意义的差异?
主要发现
- 在零假设下,受限似然比统计量的渐近分布为加权卡方分布的和,且权重可从数据中一致估计。
- 与单次分割相比,多分割方法显著降低了p值的变异性,如每组比较中500个p值的直方图所示。
- 在癌细胞系的差异回归分析中,多分割方法得到的p值为0.022(皮肤 vs. 血液),而置换检验得到的p值为0.220,表明其功效更高。
- 在肺癌与结肠癌的差异网络分析中,两种方法的p值均接近零,多分割方法得到的p值小于$10^{-4}$。
- 在合并数据上的回测显示p值接近1,证实了该方法在零假设下的有效性与稳健性。
- 该方法成功检测到了TCGA和CCLE真实高通量癌症基因组数据中的差异基因表达和网络结构。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。