Skip to main content
QUICK REVIEW

[Paper Review] Age of an allele and gene genealogies of nested subsamples for populations admitting large offspring numbers

Bjarki Eldon|arXiv (Cornell University)|Dec 8, 2012
Evolution and Genetic Dynamics17 references3 citations
TL;DR

This paper develops coalescent models for populations with large offspring number variance, deriving genealogies and allele age distributions under modified Moran models. It shows that high variance in offspring number accelerates allele turnover, resulting in significantly younger alleles and reduced genetic diversity compared to standard Kingman coalescent predictions.

ABSTRACT

Coalescent processes, including mutation, are derived from Moran type population models admitting large offspring numbers. Including mutation in the coalescent process allows for quantifying the turnover of alleles by computing the distribution of the number of original alleles still segregating in the population at a given time in the past. The turnover of alleles is considered for specific classes of the Moran model admitting large offspring numbers. Versions of the Kingman coalescent are also derived whose rates are functions of the mean and variance of the offspring distribution. High variance in the offspring distribution results in higher turnover and younger age of alleles than predicted by the usual Kingman coalescent.

Motivation & Objective

  • To model genealogical processes in populations with high variance in offspring number, which deviate from classical population genetics assumptions.
  • To extend the Moran model to incorporate random, heavy-tailed offspring distributions, enabling realistic modeling of highly fecund species.
  • To quantify how variance in offspring number affects the age and turnover rate of alleles in a population.
  • To derive coalescent processes—including mutation—that allow for recursive computation of allele ancestry and sampling probabilities in nested subsamples.
  • To compare results under multiple merger coalescent processes (evident in large offspring number models) with classical Kingman coalescent predictions.

Proposed method

  • Formalizes two generalized Moran model variants: Model 1 uses a mixture of a general offspring distribution X and a scaled uniform Y; Model 2 uses a Poisson-distributed offspring number with a mixture mean M.
  • Derives the ancestral process under these models, showing convergence to a $Λ$-coalescent with multiple mergers when offspring variance is high.
  • Incorporates neutral mutation via the infinite alleles model, enabling computation of allele age and lineage sharing probabilities.
  • Uses recursive computation to derive the distribution of the number of representatives of the oldest allele in a sample, extending results from Saunders et al. (1984).
  • Derives the probability that a nested subsample shares the most recent common ancestor with the full population, using ancestral processes with mutation.
  • Scales time in units of $N^2$ timesteps and defines a population size scaling constant $\beta = 1/c^2$, linking model parameters to effective population size.

Experimental results

Research questions

  • RQ1How does high variance in offspring number affect the age distribution of alleles in a population?
  • RQ2What is the probability that a nested subsample contains the most recent common ancestor of the entire population under large offspring number models?
  • RQ3How does the expected number of representatives of the oldest allele in a sample differ between multiple merger coalescent processes and the Kingman coalescent?
  • RQ4To what extent do the genealogies of nested subsamples deviate from Kingman coalescent predictions when offspring distributions have heavy tails?
  • RQ5How does the inclusion of mutation in the ancestral process affect the turnover rate and age of alleles in populations with large offspring variance?

Key findings

  • Populations with high variance in offspring number exhibit significantly younger alleles than predicted by the standard Kingman coalescent due to accelerated turnover.
  • The expected number of representatives of the oldest allele in a sample of size $i$, $\mathbb{E}[F_i]$, is smaller under multiple merger coalescents than under the Kingman coalescent, especially for beta coalescents.
  • For model 1 or 2, when large offspring number events are negligible, $\mathbb{E}[F_i] = \frac{i\beta + \theta}{\beta + \theta}$, which reduces to the classical result from Kelly (1977) and Saunders et al. (1984).
  • As $\psi$ decreases (indicating higher probability of large offspring events), $\mathbb{E}[F_i]$ converges to 1, indicating that the oldest allele is rarely sampled in large samples.
  • The variance of $F_i$ is given by $\mathrm{Var}[F_i] = \frac{\beta(\theta + \beta)\theta(i-1)}{(\beta + \theta)^2(2\beta + \theta)}$, showing increased sampling uncertainty under high variance.
  • Recursive computation of $F_i$ is feasible under the $\Lambda$-coalescent, but limits applicability to small sample sizes, highlighting a computational challenge in large-scale inference.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.