Skip to main content
QUICK REVIEW

[Paper Review] Two-Sample Testing in High-Dimensional Models

Nicolas Städler, Sach Mukherjee|arXiv (Cornell University)|Oct 16, 2012
Statistical Methods and Inference20 references12 citations
TL;DR

This paper proposes a general, data-splitting-based method for high-dimensional two-sample testing across models like regression and Gaussian graphical models. By combining sample splitting, $μat$-penalized estimation, and p-value aggregation over multiple splits, it achieves asymptotically valid inference via a weighted chi-square null distribution, enabling reliable testing even when $p \gg n$. The method outperforms permutation tests in power and avoids the 'p-value lottery' of single splits.

ABSTRACT

We propose novel methodology for testing equality of model parameters between two high-dimensional populations. The technique is very general and applicable to a wide range of models. The method is based on sample splitting: the data is split into two parts; on the first part we reduce the dimensionality of the model to a manageable size; on the second part we perform significance testing (p-value calculation) based on a restricted likelihood ratio statistic. Assuming that both populations arise from the same distribution, we show that the restricted likelihood ratio statistic is asymptotically distributed as a weighted sum of chi-squares with weights which can be efficiently estimated from the data. In high-dimensional problems, a single data split can result in a "p-value lottery". To ameliorate this effect, we iterate the splitting process and aggregate the resulting p-values. This multi-split approach provides improved p-values. We illustrate the use of our general approach in two-sample comparisons of high-dimensional regression models ("differential regression") and graphical models ("differential network"). In both cases we show results on simulated data as well as real data from recent, high-throughput cancer studies.

Motivation & Objective

  • To address the lack of reliable significance testing in high-dimensional two-sample problems where $p \gg n$.
  • To develop a general framework applicable to diverse models, including high-dimensional regression and Gaussian graphical models.
  • To overcome the instability of single data splits ('p-value lottery') through iterative splitting and p-value aggregation.
  • To provide a theoretically grounded approach with asymptotic null distribution as a weighted sum of chi-squares, estimable from data.
  • To enable statistical testing of differential relationships in high-throughput biological data, such as differential networks in cancer genomics.

Proposed method

  • Split data from both populations into two parts: one for model screening, one for testing.
  • Use $\ell_1$-penalized maximum likelihood estimation on the first split to select active parameter sets, ensuring sparsity and screening properties.
  • On the second split, compute a restricted likelihood ratio statistic comparing non-nested models (individual vs. pooled parameter spaces).
  • Establish the asymptotic null distribution of the restricted likelihood ratio as a weighted sum of independent $\chi^2$ variables, with weights estimated from the data.
  • Iterate the splitting process multiple times and aggregate p-values to reduce variability and improve reliability.
  • Apply the method to differential regression (high-dimensional linear models) and differential network (Gaussian graphical models) via penalized likelihood estimation.

Experimental results

Research questions

  • RQ1Can a general two-sample testing framework be developed for high-dimensional models where $p \gg n$?
  • RQ2How can significance testing be reliably performed when classical likelihood ratio tests break down in high dimensions?
  • RQ3Can data splitting and p-value aggregation mitigate the instability of single-split p-values in high-dimensional settings?
  • RQ4What is the asymptotic null distribution of the restricted likelihood ratio statistic in non-nested, high-dimensional models?
  • RQ5Can this method detect biologically meaningful differences in gene regulatory networks or regression effects across cancer subtypes?

Key findings

  • The restricted likelihood ratio statistic asymptotically follows a weighted sum of chi-squared distributions under the null, with weights that can be consistently estimated from the data.
  • The multi-split approach significantly reduces p-value variability compared to single splits, as shown by histograms of 500 p-values per comparison.
  • In differential regression on cancer cell lines, the multi-split method yielded a p-value of 0.022 for skin vs. haem, while the permutation test gave 0.220, indicating higher power.
  • For differential network analysis between lung and colon cancers, both methods produced p-values near zero, with the multi-split method yielding $<10^{-4}$.
  • Back-testing on pooled data showed p-values near one, confirming the method's validity and robustness under the null.
  • The method successfully detected differential gene expression and network structure in real high-throughput cancer genomics data from TCGA and CCLE.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.