Skip to main content
QUICK REVIEW

[Paper Review] A mixture model for determining SARS-Cov-2 variant composition in pooled samples

Silva, Israel Tojal da|arXiv (Cornell University)|Jan 1, 2022
SARS-CoV-2 detection and testing33 references54 citations
TL;DR

This paper proposes a statistical mixture model to estimate the relative frequencies of SARS-CoV-2 variants in pooled samples using genomic polymorphisms (SNPs/indels) as markers. The method uses maximum likelihood estimation on sequencing read counts at variant-specific loci, accounting for latent variant contributions and coverage variation. It accurately recovers variant proportions in simulations and shows strong correlation with clinical epidemiological trends in real wastewater data from Switzerland, demonstrating its utility for population-level viral surveillance.

ABSTRACT

Despite of the fast development of highly effective vaccines to control the current COVID$-$19 pandemic, the unequal distribution and availability of these vaccines worldwide and the number of people infected in the world lead to the continuous emergence of SARS-CoV-2 (Severe Acute Respiratory Syndrome coronavirus 2) variants of concern. It is likely that real-time genomic surveillance will be continuously needed as an unceasing monitoring tool, necessary to follow the spillover of the disease spread and the evolution of the virus. In this context, new genomic variants of SARS-CoV-2 that may emerge as a response to selective pressure, including variants refractory to current vaccines, makes genomic surveillance programs tools of utmost importance. Here propose a statistical model for the estimation of the relative frequencies of SARS-CoV-2 variants in pooled samples. This model is built by considering a previously defined selection of genomic polymorphisms that characterize SARS-CoV-2 variants. The methods described here support both raw sequencing reads for polymorphisms-based markers calling and predefined markers in the VCF format. Results obtained by using simulated data show that our method is quite effective in recovering the correct variant proportions. Further, results obtained by considering longitudinal data from wastewater samples of two locations in Switzerland agree well with those describing the epidemiological evolution of COVID-19 variants in clinical samples of these locations. Our results show that the described method can be a valuable tool for tracking the proportions of SARS-CoV-2 variants.

Motivation & Objective

  • To develop a statistical method for estimating relative variant frequencies in pooled SARS-CoV-2 samples, where multiple viral lineages co-occur.
  • To address the challenge of inferring variant composition from short-read sequencing data with low and variable coverage.
  • To enable cost-effective, high-throughput surveillance by leveraging pooled samples such as wastewater.
  • To validate the method using both simulated data and real-world wastewater sequencing from Switzerland.
  • To support public health monitoring by aligning variant frequency trends with clinical epidemiological data.

Proposed method

  • The method models variant composition using a mixture model based on preselected genomic polymorphisms (SNPs/indels) that define SARS-CoV-2 variants.
  • It employs a likelihood function that integrates over latent variables representing the contribution of each variant to observed reference and alternate allele counts at each polymorphic site.
  • The likelihood is computed under the assumption of multinomial distribution of reads across variants, with variant proportions wj as parameters to be estimated.
  • Maximum likelihood estimation is used to infer the vector w = (w1, ..., wv) of relative variant frequencies, subject to ∑wj = 1 and 0 ≤ wj ≤ 1.
  • The model accounts for variable coverage across loci by incorporating total read counts ti = ca_i + cr_i at each polymorphic site.
  • A bootstrap resampling approach is applied to estimate standard errors and quantify uncertainty in variant frequency predictions.

Experimental results

Research questions

  • RQ1Can a statistical mixture model accurately estimate the relative frequencies of SARS-CoV-2 variants in pooled clinical or environmental samples?
  • RQ2How well does the model perform under low sequencing depth and variable coverage conditions?
  • RQ3To what extent do variant frequency estimates from wastewater samples correlate with clinical epidemiological trends?
  • RQ4Can the method detect emerging variants and track lineage dynamics in real-world settings?
  • RQ5How robust is the method to intra-variant genetic heterogeneity and shared polymorphisms across lineages?

Key findings

  • The method accurately recovers true variant proportions in simulated data, even at low sequencing depths, demonstrating high accuracy and sensitivity.
  • In wastewater samples from Lausanne and Zurich, the estimated variant frequencies closely matched the epidemiological trends observed in clinical sequencing data from the same regions.
  • The longitudinal analysis of 122 wastewater samples showed consistent detection of the Alpha variant's rise and decline, aligning with regional clinical data.
  • The model's performance was robust to variable coverage and stochastic sampling, with bootstrap-based uncertainty estimates providing reliable prediction intervals.
  • The method successfully identified variant transitions and dominance shifts in the viral population over time, confirming its utility for surveillance.
  • The software pipeline is scalable, supports custom marker sets in VCF format, and enables parallelized, efficient processing via a workflow engine.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.