Skip to main content
QUICK REVIEW

[论文解读] A mixture model for determining SARS-Cov-2 variant composition in pooled samples

Silva, Israel Tojal da|arXiv (Cornell University)|Jan 1, 2022
SARS-CoV-2 detection and testing参考文献 33被引用 54
一句话总结

本文提出了一种统计混合模型,利用基因组多态性(SNPs/indels)作为标记,通过在变异特异性位点的测序读数进行最大似然估计,估算混合样本中SARS-CoV-2变异株的相对频率。该方法考虑了潜在变异株的贡献和覆盖度变化,能在模拟数据中准确恢复变异株比例,并在瑞士真实污水数据中与临床流行病学趋势表现出强相关性,证明了其在人群水平病毒监测中的实用性。

ABSTRACT

Despite of the fast development of highly effective vaccines to control the current COVID$-$19 pandemic, the unequal distribution and availability of these vaccines worldwide and the number of people infected in the world lead to the continuous emergence of SARS-CoV-2 (Severe Acute Respiratory Syndrome coronavirus 2) variants of concern. It is likely that real-time genomic surveillance will be continuously needed as an unceasing monitoring tool, necessary to follow the spillover of the disease spread and the evolution of the virus. In this context, new genomic variants of SARS-CoV-2 that may emerge as a response to selective pressure, including variants refractory to current vaccines, makes genomic surveillance programs tools of utmost importance. Here propose a statistical model for the estimation of the relative frequencies of SARS-CoV-2 variants in pooled samples. This model is built by considering a previously defined selection of genomic polymorphisms that characterize SARS-CoV-2 variants. The methods described here support both raw sequencing reads for polymorphisms-based markers calling and predefined markers in the VCF format. Results obtained by using simulated data show that our method is quite effective in recovering the correct variant proportions. Further, results obtained by considering longitudinal data from wastewater samples of two locations in Switzerland agree well with those describing the epidemiological evolution of COVID-19 variants in clinical samples of these locations. Our results show that the described method can be a valuable tool for tracking the proportions of SARS-CoV-2 variants.

研究动机与目标

  • 开发一种统计方法,用于估算混合SARS-CoV-2样本中相对变异株频率,其中多个病毒谱系共存。
  • 解决在低且可变覆盖度的短读长测序数据中推断变异株组成的问题。
  • 通过利用如污水等混合样本,实现成本效益高、高通量的监测。
  • 使用模拟数据和来自瑞士的真实污水测序数据验证该方法。
  • 通过将变异株频率趋势与临床流行病学数据对齐,支持公共卫生监测。

提出的方法

  • 该方法基于预先选定的基因组多态性(SNPs/indels)构建混合模型,用于建模变异株组成,这些多态性定义了SARS-CoV-2变异株。
  • 采用似然函数,对每个多态性位点上代表各变异株对观察到的参考等位基因和替代等位基因计数贡献的隐变量进行积分。
  • 在读数在各变异株之间呈多项分布的假设下计算似然函数,其中变异株比例wj作为待估计的参数。
  • 使用最大似然估计推断相对变异株频率向量w = (w1, ..., wv),受约束于∑wj = 1且0 ≤ wj ≤ 1。
  • 通过在每个多态性位点引入总读数ti = ca_i + cr_i,考虑各位点间覆盖度的差异。
  • 采用自助重采样方法估计标准误,并量化变异株频率预测中的不确定性。

实验结果

研究问题

  • RQ1统计混合模型能否准确估算混合临床或环境样本中SARS-CoV-2变异株的相对频率?
  • RQ2在测序深度低且覆盖度可变的条件下,该模型表现如何?
  • RQ3污水样本中的变异株频率估计与临床流行病学趋势的相关性有多大?
  • RQ4该方法能否在真实世界环境中检测到新发变异株并追踪谱系动态?
  • RQ5该方法对变异株内遗传异质性和谱系间共享多态性有多大的鲁棒性?

主要发现

  • 在模拟数据中,即使测序深度较低,该方法也能准确恢复真实的变异株比例,表现出高准确性和敏感性。
  • 在洛桑和苏黎世的污水样本中,估算的变异株频率与同一地区临床测序数据中观察到的流行病学趋势高度一致。
  • 对122份污水样本的纵向分析显示,该模型一致检测到了Alpha变异株的上升和下降趋势,与区域临床数据相符。
  • 该模型对可变覆盖度和随机抽样具有鲁棒性,基于自助法的不确定性估计提供了可靠的预测区间。
  • 该方法成功识别了病毒群体随时间推移的变异株转换和主导性转变,证实了其在监测中的实用性。
  • 该软件工作流可扩展,支持VCF格式的自定义标记集,并可通过工作流引擎实现并行化和高效处理。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。