Skip to main content
QUICK REVIEW

[Paper Review] Set-Based Tests for Genetic Association Using the Generalized Berk-Jones Statistic

Ryan Sun, Xihong Lin|arXiv (Cornell University)|Oct 6, 2017
Genetic Associations and Epidemiology24 references15 citations
TL;DR

This paper proposes the Generalized Berk-Jones (GBJ) test for set-based genetic association analysis, which extends the Berk-Jones statistic to account for correlation among SNPs in a set, improving power under moderate sparsity. The method enables analytic p-value calculation using summary statistics and demonstrates robust performance across varying signal sparsity and linkage disequilibrium patterns, outperforming existing methods like GHC and MinP in many scenarios.

ABSTRACT

Studying the effects of groups of Single Nucleotide Polymorphisms (SNPs), as in a gene, genetic pathway, or network, can provide novel insight into complex diseases, above that which can be gleaned from studying SNPs individually. Common challenges in set-based genetic association testing include weak effect sizes, correlation between SNPs in a SNP-set, and scarcity of signals, with single-SNP effects often ranging from extremely sparse to moderately sparse in number. Motivated by these challenges, we propose the Generalized Berk-Jones (GBJ) test for the association between a SNP-set and outcome. The GBJ extends the Berk-Jones (BJ) statistic by accounting for correlation among SNPs, and it provides advantages over the Generalized Higher Criticism (GHC) test when signals in a SNP-set are moderately sparse. We also provide an analytic p-value calculation procedure for SNP-sets of any finite size. Using this p-value calculation, we illustrate that the rejection region for GBJ can be described as a compromise of those for BJ and GHC. We develop an omnibus statistic as well, and we show that this omnibus test is robust to the degree of signal sparsity. An additional advantage of our method is the ability to conduct inference using individual SNP summary statistics from a genome-wide association study. We evaluate the finite sample performance of the GBJ though simulation studies and application to gene-level association analysis of breast cancer risk.

Motivation & Objective

  • To address the challenge of low power in set-based genetic association testing when signals are sparse and SNPs are correlated due to linkage disequilibrium.
  • To extend the Berk-Jones statistic to account for correlation among SNPs in a SNP-set, improving finite-sample performance.
  • To develop an analytic p-value calculation for GBJ applicable to SNP-sets of any finite size, enabling inference from summary statistics.
  • To create an omnibus test that maintains robust power across diverse signal sparsity levels and correlation structures.
  • To provide a method that outperforms existing approaches like GHC, MinP, and SKAT in moderately sparse settings common in gene- and pathway-level GWAS.

Proposed method

  • The GBJ test modifies the standard Berk-Jones statistic to incorporate a correlation matrix among SNPs, adjusting for LD-induced dependence.
  • It uses a supremum-based test statistic that maximizes evidence across all possible subsets of SNPs in a set, enhancing sensitivity to sparse signals.
  • An analytic p-value calculation is derived using a generalized chi-square approximation, enabling exact inference without permutation.
  • The method is generalized to a class of supremum-based global tests, allowing valid inference for HC, GHC, BJ, and similar statistics under correlation.
  • An omnibus test is constructed by combining GBJ with other statistics, ensuring robustness to unknown signal sparsity and correlation patterns.
  • The approach leverages individual SNP summary statistics from GWAS, enabling application without access to raw genotype data.

Experimental results

Research questions

  • RQ1Can the Berk-Jones statistic be adapted to maintain high power under correlated SNPs in SNP-sets?
  • RQ2How does the performance of GBJ compare to GHC and MinP in moderately sparse signal regimes common in gene-level GWAS?
  • RQ3Can an analytic p-value be derived for the generalized Berk-Jones test under arbitrary correlation structures?
  • RQ4Does the omnibus test based on GBJ maintain high power across diverse signal sparsity and correlation levels?
  • RQ5Can the rejection region of GBJ be interpreted as a compromise between those of BJ and GHC, balancing sensitivity across signal locations?

Key findings

  • GBJ outperforms the standard Berk-Jones test in high linkage disequilibrium settings, where the original BJ loses power due to unaccounted correlation.
  • The rejection region of GBJ closely approximates the optimal boundary at both the most extreme and moderately extreme order statistics, balancing sensitivity across signal locations.
  • In simulations, GBJ achieves higher power than GHC and MinP under moderate sparsity, and outperforms SKAT when signals are sparse and uncorrelated.
  • The omnibus test based on GBJ maintains robust power across all sparsity levels and is never the worst-performing method, though it is rarely the best.
  • Application to the CGEMS breast cancer GWAS data shows GBJ produces the most significant p-values among tested methods, indicating strong empirical performance.
  • The analytic p-value calculation for GBJ is accurate and computationally efficient, enabling widespread use in summary-Statistic-based GWAS.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.