Skip to main content
QUICK REVIEW

[Paper Review] Genetic similarity versus genetic ancestry groups as sample descriptors in human genetics

Graham Coop|arXiv (Cornell University)|Jul 23, 2022
Genetic Associations and Epidemiology21 citations
TL;DR

The paper argues that genetic ancestry group labels are imprecise and often misleading, and recommends using explicit statements of genetic similarity/relatedness to describe samples instead.

ABSTRACT

A common sample descriptor in human genomics studies is that of 'genetic ancestry group', with terms such as 'European genetic ancestry' or 'East Asian genetic ancestry' frequently used in publications to describe the genetics of groups of individuals based on the analysis of their genotypes. In this Perspective, I argue that these terms are imprecise and potentially misleading and that, for most applications, simple statements of genetic similarity represent a more accurate description.

Motivation & Objective

  • Motivate why genetic ancestry group labels are imprecise and potentially misleading.
  • Explain how genetic similarity/relatedness provides a more accurate and communicative sample descriptor.
  • Discuss practical implications for data subsetting, analysis, and interpretation in human genetics.

Proposed method

  • Review conceptual foundations of genetic variation, ancestry, and genetic similarity.
  • Analyze how common ancestry-labeling methods (e.g., PCA clustering, STRUCTURE/ADMIXTURE, haplotype-based local ancestry) function as descriptors.
  • Argue how these labels effectively convey genetic similarity rather than defined ancestral groups.
  • Propose terminology shifts toward describing genetic similarity between samples (e.g., 'genetically similar to XX samples on PC axes').
  • Discuss limitations, potential timeframes, and alternatives for describing ancestry across time epochs.

Experimental results

Research questions

  • RQ1When and why are genetic ancestry group labels used in human genetics, and what do they actually describe?
  • RQ2What are the limitations and risks of using ancestry-based labels for describing genetic data and health associations?
  • RQ3How can researchers reframe sample descriptors to emphasize genetic similarity and relatedness instead of ancestral categories?
  • RQ4What practical impact would a shift to similarity-based descriptors have on data subsetting, analyses, and communication across studies?

Key findings

  • Genetic ancestry labels are proxy statements of genetic similarity to reference panels, not precise descriptions of defined ancestral groups.
  • Ancestry labels can mislead by implying homogeneous groups and by conflating genetics with social/environmental factors.
  • There is continuous, not discrete, genetic variation; ancestry labels often depend on reference panels and chosen timeframes, causing instability across studies.
  • Researchers frequently subset data for methodological reasons, making clear similarity-based descriptors more accurate and informative for matching and controls.
  • A shift toward 'genetic similarity/relatedness' descriptors improves clarity and reduces baggage associated with ancestry language.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.