Skip to main content
QUICK REVIEW

[Paper Review] On the genetic architecture of intelligence and other quantitative traits

Stephen Hsu|arXiv (Cornell University)|Aug 14, 2014
Genetic Associations and Epidemiology3 references3 citations
TL;DR

This paper investigates the genetic architecture of intelligence and other quantitative traits like height, proposing that a large number of common and rare genetic variants—particularly around 10,000 moderately rare variants of mostly negative effect—contribute to normal population variation. Using compressed sensing and L1-penalized regression, the author estimates that characterizing these traits requires sample sizes on the order of 100× the number of causal loci, or ~1 million individuals, to achieve statistical power for predicting g (general intelligence factor) from genotype.

ABSTRACT

How do genes affect cognitive ability or other human quantitative traits such as height or disease risk? Progress on this challenging question is likely to be significant in the near future. I begin with a brief review of psychometric measurements of intelligence, introducing the idea of a "general factor" or g score. The main results concern the stability, validity (predictive power), and heritability of adult g. The largest component of genetic variance for both height and intelligence is additive (linear), leading to important simplifications in predictive modeling and statistical estimation. Due mainly to the rapidly decreasing cost of genotyping, it is possible that within the coming decade researchers will identify loci which account for a significant fraction of total g variation. In the case of height analogous efforts are well under way. I describe some unpublished results concerning the genetic architecture of height and cognitive ability, which suggest that roughly 10k moderately rare causal variants of mostly negative effect are responsible for normal population variation. Using results from Compressed Sensing (L1-penalized regression), I estimate the statistical power required to characterize both linear and nonlinear models for quantitative traits. The main unknown parameter s (sparsity) is the number of loci which account for the bulk of the genetic variation. The required sample size is of order 100s, or roughly a million in the case of cognitive ability.

Motivation & Objective

  • To understand the genetic basis of complex human traits such as intelligence, height, and disease risk, focusing on polygenic architecture.
  • To assess the feasibility of predicting cognitive ability from genotype using emerging genomic technologies.
  • To estimate the sample size required to identify and model the genetic loci underlying quantitative traits like intelligence and height.
  • To evaluate the role of additive genetic variance and the potential for nonlinear models in explaining trait variation.
  • To explore the implications of genomic prediction for embryo selection, cognitive enhancement, and public health.

Proposed method

  • Uses compressed sensing (L1-penalized regression) to model high-dimensional genetic data and estimate the number of causal loci.
  • Applies statistical estimation techniques to determine the minimum sample size required for reliable detection of genetic variants.
  • Analyzes existing twin and adoption studies to estimate heritability of g (general factor of intelligence).
  • Models genetic architecture assuming a large number of moderately rare variants with small to moderate effects.
  • Estimates that sample size should scale as ~100× the number of causal loci (s) for effective detection.
  • Relies on data from SNP genotyping and whole-genome sequencing, with costs declining rapidly to ~$100 and ~$1000 per individual, respectively.

Experimental results

Research questions

  • RQ1What is the genetic architecture of intelligence, and how many loci contribute to variation in the general factor g?
  • RQ2How many genetic variants are needed to explain a significant fraction of the variance in cognitive ability and height?
  • RQ3What sample size is required to reliably detect and model the genetic variants underlying complex quantitative traits?
  • RQ4To what extent is the genetic variance in intelligence and height additive, and how does this simplify predictive modeling?
  • RQ5Can compressed sensing methods effectively recover the genetic basis of polygenic traits from high-dimensional genomic data?

Key findings

  • The largest component of genetic variance for both intelligence and height is additive, simplifying predictive modeling and statistical estimation.
  • The genetic architecture of intelligence and height is likely shaped by approximately 10,000 moderately rare causal variants, mostly with negative effects on trait variation.
  • Sample size requirements for detecting causal loci scale as ~100× the number of loci (s), implying a need for roughly one million individuals to characterize the genetic basis of intelligence.
  • Rapidly declining genotyping costs suggest that identifying loci accounting for a significant fraction of g variation may be feasible within the next decade.
  • Heritability of g is estimated at 50–80%, based on twin and adoption studies, with strong predictive power for outcomes such as education, income, and longevity.
  • The results imply that embryo selection based on polygenic scores may become technically feasible and ethically consequential within a short timeframe.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.