Skip to main content
QUICK REVIEW

[Paper Review] Improving power of genetic association studies by extreme phenotype sampling: a review and some new results

Thea Bjørnland, Anja Bye|arXiv (Cornell University)|Jan 5, 2017
Genetic and phenotypic traits in livestockBiochemistry, Genetics and Molecular Biology6 citations
TL;DR

This paper proposes likelihood-based inference methods for genetic association studies using extreme phenotype sampling, where only individuals with the most extreme trait values are genotyped. It demonstrates that this design can significantly improve statistical power over random sampling, particularly for detecting common variant effects and gene-environment interactions, though performance varies by sampling strategy and trait distribution.

ABSTRACT

Extreme phenotype sampling is a selective genotyping design for genetic association studies where only individuals with extreme values of a continuous trait are genotyped for a set of genetic variants. Under financial or other limitations, this design is assumed to improve the power to detect associations between genetic variants and the trait, compared to randomly selecting the same number of individuals for genotyping. Here we present extensions of likelihood models that can be used for inference when the data are sampled according to the extreme phenotype sampling design. Computational methods for parameter estimation and hypothesis testing are provided. We consider methods for common variant genetic effects and gene-environment interaction effects in linear regression models with a normally distributed trait. We use simulated and real data to show that extreme phenotype sampling can be powerful compared to random sampling, but that this does not hold for all extreme sampling methods and situations.

Motivation & Objective

  • To develop robust statistical methods for inference in genetic association studies when only extreme phenotype individuals are genotyped.
  • To evaluate whether extreme phenotype sampling increases power compared to random sampling under realistic genetic and trait assumptions.
  • To extend likelihood-based models to accommodate both common variant effects and gene-environment interactions in normally distributed traits.
  • To provide computational tools for parameter estimation and hypothesis testing under the extreme phenotype sampling design.
  • To assess the conditions under which extreme phenotype sampling outperforms random sampling in terms of statistical power.

Proposed method

  • Develops likelihood models tailored to extreme phenotype sampling designs, accounting for the selection mechanism in parameter estimation.
  • Applies likelihood ratio tests and score tests to assess genetic effects and gene-environment interactions under the sampling scheme.
  • Uses numerical optimization techniques to estimate model parameters efficiently despite the truncated sampling distribution.
  • Incorporates the normality assumption of the trait into the likelihood function to maintain statistical efficiency.
  • Adapts standard linear regression models to include interaction terms between genetic variants and environmental factors under the extreme sampling framework.
  • Validates the method using both simulated datasets with known genetic effects and real-world genetic data to assess empirical performance.

Experimental results

Research questions

  • RQ1Does extreme phenotype sampling increase statistical power for detecting genetic associations compared to random sampling?
  • RQ2How do different sampling strategies (e.g., top/bottom 10%, 20%) affect the power and accuracy of association tests?
  • RQ3Can likelihood-based inference methods maintain valid Type I error rates and power when applied to extreme phenotype samples?
  • RQ4How does gene-environment interaction detection performance compare between extreme and random sampling designs?
  • RQ5Under what conditions does extreme phenotype sampling fail to improve power despite theoretical expectations?

Key findings

  • Extreme phenotype sampling can significantly increase statistical power for detecting common variant effects compared to random sampling, particularly when the trait is normally distributed.
  • The power gain is most substantial when the extreme sampling fraction is moderate (e.g., top and bottom 10–20% of the trait distribution).
  • Likelihood-based inference methods maintain appropriate Type I error rates and provide valid statistical inference under the extreme sampling design.
  • Gene-environment interaction effects can be detected with higher power using extreme phenotype sampling, especially when the interaction effect size is moderate.
  • However, power gains diminish or disappear when the trait distribution deviates from normality or when sampling fractions are too narrow or too broad.
  • The method's performance is sensitive to the choice of sampling threshold and trait variance structure, highlighting the need for careful design planning.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.