Skip to main content
QUICK REVIEW

[Paper Review] Gains in Power from Structured Two-Sample Tests of Means on Graphs

Laurent Jacob, Pierre Neuvial|Collection of Biostatistics Research Archive|Sep 27, 2010
Bioinformatics and Genomic Networks41 references20 citations
TL;DR

This paper proposes a graph-structured two-sample test of means that leverages known network topology to boost statistical power when detecting differential expression in genes or other variables. By projecting test statistics onto low-frequency graph-Fourier modes—assuming smooth mean shifts on the graph—it achieves higher power than classical tests like Hotelling’s $T^2$, especially in high-dimensional, structured data such as KEGG pathways in cancer genomics.

ABSTRACT

We consider multivariate two-sample tests of means, where the location shift between the two populations is expected to be related to a known graph structure. An important application of such tests is the detection of differentially expressed genes between two patient populations, as shifts in expression levels are expected to be coherent with the structure of graphs reflecting gene properties such as biological process, molecular function, regulation, or metabolism. For a fixed graph of interest, we demonstrate that accounting for graph structure can yield more powerful tests under the assumption of smooth distribution shift on the graph. We also investigate the identification of non-homogeneous subgraphs of a given large graph, which poses both computational and multiple testing problems. The relevance and benefits of the proposed approach are illustrated on synthetic data and on breast cancer gene expression data analyzed in context of KEGG pathways.

Motivation & Objective

  • Address the limitation of classical multivariate two-sample tests (e.g., Hotelling’s $T^2$) that ignore known biological or structural relationships among variables.
  • Improve detection power in differential expression analysis by incorporating prior knowledge of gene networks (e.g., KEGG pathways) as graph structures.
  • Develop efficient algorithms to identify non-homogeneous subgraphs within a large graph, addressing computational and multiple testing challenges.
  • Enable systematic discovery of biologically relevant, coherent expression shifts across gene networks, improving interpretability and biological insight.

Proposed method

  • Model the mean shift $\delta = \mu_1 - \mu_2$ as a function on a known undirected graph $\mathcal{G} = (\mathcal{V}, \mathcal{E})$, assuming smoothness via graph Laplacian spectral decomposition.
  • Transform the test statistics into the graph-Fourier basis, focusing on low-frequency components that correspond to smooth shifts on the graph.
  • Construct a test statistic based on the sum of squared coefficients of the first few graph-Fourier modes, which captures smooth deviations from the null hypothesis.
  • Use branch-and-bound pruning algorithms to efficiently search for non-homogeneous subgraphs, reducing the number of explicit tests while maintaining power.
  • Apply permutation-based procedures to estimate the distribution of false positives under the null, accounting for dependence among subgraph tests.
  • Integrate regularization techniques (e.g., from Tai and Speed, 2008) to improve numerical stability when estimating covariance matrices in high dimensions.

Experimental results

Research questions

  • RQ1Can incorporating graph structure into two-sample tests of means lead to higher statistical power under smooth mean shift alternatives?
  • RQ2How can one efficiently identify non-homogeneous subgraphs within a large network without exhaustive testing of all possible subgraphs?
  • RQ3What is the impact of graph structure on the detection of differentially expressed genes in real-world genomic data, such as in KEGG pathways?
  • RQ4How do the proposed methods compare to classical multivariate tests (e.g., Hotelling’s $T^2$) and univariate gene-level tests in terms of power and biological interpretability?
  • RQ5Can the method detect biologically meaningful, coherent expression changes even when individual genes show weak or inconsistent signals?

Key findings

  • The graph-structured test achieves significant power gains over Hotelling’s $T^2$ under smooth-shift alternatives, particularly when the true mean shift aligns with the graph structure.
  • On synthetic data, the proposed method detects smooth shifts with higher power than classical tests, especially in high-dimensional settings.
  • In the breast cancer microarray dataset from Loi et al. (2008), the method successfully identified two overlapping subgraphs in KEGG pathways associated with tamoxifen resistance, with coherent expression patterns.
  • The branch-and-bound algorithm reduces the number of explicitly tested subgraphs while maintaining high detection power, enabling scalable exploration of large networks.
  • Permutation-based false positive control is effective despite unknown and highly dependent test counts, providing valid inference in subgraph discovery.
  • The method outperforms standard univariate and multivariate approaches by leveraging biological network topology, revealing biologically relevant gene sets that are missed otherwise.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.