[Paper Review] Two-Sample Testing in High-Dimensional Models
This paper proposes a general, data-splitting-based method for high-dimensional two-sample testing across models like regression and Gaussian graphical models. By combining sample splitting, $μat$-penalized estimation, and p-value aggregation over multiple splits, it achieves asymptotically valid inference via a weighted chi-square null distribution, enabling reliable testing even when $p \gg n$. The method outperforms permutation tests in power and avoids the 'p-value lottery' of single splits.
We propose novel methodology for testing equality of model parameters between two high-dimensional populations. The technique is very general and applicable to a wide range of models. The method is based on sample splitting: the data is split into two parts; on the first part we reduce the dimensionality of the model to a manageable size; on the second part we perform significance testing (p-value calculation) based on a restricted likelihood ratio statistic. Assuming that both populations arise from the same distribution, we show that the restricted likelihood ratio statistic is asymptotically distributed as a weighted sum of chi-squares with weights which can be efficiently estimated from the data. In high-dimensional problems, a single data split can result in a "p-value lottery". To ameliorate this effect, we iterate the splitting process and aggregate the resulting p-values. This multi-split approach provides improved p-values. We illustrate the use of our general approach in two-sample comparisons of high-dimensional regression models ("differential regression") and graphical models ("differential network"). In both cases we show results on simulated data as well as real data from recent, high-throughput cancer studies.
Motivation & Objective
- To address the lack of reliable significance testing in high-dimensional two-sample problems where $p \gg n$.
- To develop a general framework applicable to diverse models, including high-dimensional regression and Gaussian graphical models.
- To overcome the instability of single data splits ('p-value lottery') through iterative splitting and p-value aggregation.
- To provide a theoretically grounded approach with asymptotic null distribution as a weighted sum of chi-squares, estimable from data.
- To enable statistical testing of differential relationships in high-throughput biological data, such as differential networks in cancer genomics.
Proposed method
- Split data from both populations into two parts: one for model screening, one for testing.
- Use $\ell_1$-penalized maximum likelihood estimation on the first split to select active parameter sets, ensuring sparsity and screening properties.
- On the second split, compute a restricted likelihood ratio statistic comparing non-nested models (individual vs. pooled parameter spaces).
- Establish the asymptotic null distribution of the restricted likelihood ratio as a weighted sum of independent $\chi^2$ variables, with weights estimated from the data.
- Iterate the splitting process multiple times and aggregate p-values to reduce variability and improve reliability.
- Apply the method to differential regression (high-dimensional linear models) and differential network (Gaussian graphical models) via penalized likelihood estimation.
Experimental results
Research questions
- RQ1Can a general two-sample testing framework be developed for high-dimensional models where $p \gg n$?
- RQ2How can significance testing be reliably performed when classical likelihood ratio tests break down in high dimensions?
- RQ3Can data splitting and p-value aggregation mitigate the instability of single-split p-values in high-dimensional settings?
- RQ4What is the asymptotic null distribution of the restricted likelihood ratio statistic in non-nested, high-dimensional models?
- RQ5Can this method detect biologically meaningful differences in gene regulatory networks or regression effects across cancer subtypes?
Key findings
- The restricted likelihood ratio statistic asymptotically follows a weighted sum of chi-squared distributions under the null, with weights that can be consistently estimated from the data.
- The multi-split approach significantly reduces p-value variability compared to single splits, as shown by histograms of 500 p-values per comparison.
- In differential regression on cancer cell lines, the multi-split method yielded a p-value of 0.022 for skin vs. haem, while the permutation test gave 0.220, indicating higher power.
- For differential network analysis between lung and colon cancers, both methods produced p-values near zero, with the multi-split method yielding $<10^{-4}$.
- Back-testing on pooled data showed p-values near one, confirming the method's validity and robustness under the null.
- The method successfully detected differential gene expression and network structure in real high-throughput cancer genomics data from TCGA and CCLE.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.