[论文解读] eDNAPlus: A unifying modelling framework for DNA-based biodiversity monitoring
eDNAPlus 提出了一种统一的分层贝叶斯模型,用于基于DNA的生物多样性监测,该模型同时考虑了宏条形码工作流程中所有主要的误差和噪声来源,实现了对种内生物量变化及其与环境协变量关联的稳健估计,同时建模了种间和站点间的相关性。该框架通过自适应MCMC和重参数化提高了推断效率,并在马莱塞陷阱数据上表现出色,揭示了海拔和距道路距离的影响,并识别出高保护价值区域。
DNA-based biodiversity surveys involve collecting physical samples from survey sites and assaying the contents in the laboratory to detect species via their diagnostic DNA sequences. DNA-based surveys are increasingly being adopted for biodiversity monitoring. The most commonly employed method is metabarcoding, which combines PCR with high-throughput DNA sequencing to amplify and then read `DNA barcode' sequences. This process generates count data indicating the number of times each DNA barcode was read. However, DNA-based data are noisy and error-prone, with several sources of variation. In this paper, we present a unifying modelling framework for DNA-based data allowing for all key sources of variation and error in the data-generating process. The model can estimate within-species biomass changes across sites and link those changes to environmental covariates, while accounting for species and sites correlation. Inference is performed using MCMC, where we employ Gibbs or Metropolis-Hastings updates with Laplace approximations. We also implement a re-parameterisation scheme, appropriate for crossed-effects models, leading to improved mixing, and an adaptive approach for updating latent variables, reducing computation time. We discuss study design and present theoretical and simulation results to guide decisions on replication at different stages and on the use of quality control methods. We demonstrate the new framework on a dataset of Malaise-trap samples. We quantify the effects of elevation and distance-to-road on each species, infer species correlations, and produce maps identifying areas of high biodiversity, which can be used to rank areas by conservation value. We estimate the level of noise between sites and within sample replicates, and the probabilities of error at the PCR stage, which are close to zero for most species considered, validating the employed laboratory processing.
研究动机与目标
- 为解决现有统计框架中缺乏对DNA基生物多样性监测工作流程中所有关键变异来源、误差和噪声的全面建模问题。
- 在考虑种间和站点间相关性的同时,估计不同站点间的种内生物量变化。
- 以统计严谨的方式将生物量变化与环境协变量(如海拔和距道路距离)相关联。
- 通过自适应MCMC、Gibbs/梅特罗波利斯-黑斯廷斯更新以及交叉效应模型的重参数化,提升推断的计算效率。
- 通过理论分析和基于模拟的洞察,为研究设计提供支持,优化各调查阶段的重复设置与质量控制策略。
提出的方法
- 该框架采用分层交叉效应模型,联合建模不同站点的物种特异性生物量、检测误差和环境协变量。
- 使用MCMC推断,结合Gibbs和梅特罗波利斯-黑斯廷斯更新,并通过拉普拉斯近似提高潜变量估计的效率。
- 应用重参数化方案以改善MCMC链中混合性,尤其针对交叉随机效应。
- 对潜变量采用自适应更新策略,在不损失准确性的前提下显著减少计算时间。
- 模型设计用于直接处理宏条形码的原始计数数据,避免转换为二元存在/缺失状态,从而规避人为阈值的影响。
- 该框架可扩展至多引物设计,并可通过变分贝叶斯或低秩矩阵近似方法适配宏基因组数据。
实验结果
研究问题
- RQ1如何在一个统一的统计框架内同时建模环境DNA宏条形码工作流程中的所有主要误差和噪声来源?
- RQ2环境协变量(如海拔和距道路距离)在多大程度上影响调查站点中物种的生物量?
- RQ3如何在考虑种间和站点间相关性的情况下,可靠估计种内生物量变化?
- RQ4PCR水平的质量控制对推断准确性和模型稳健性有何影响?
- RQ5如何优化研究设计以实现重复设置与质量控制的最优化,从而最大化统计效能并最小化偏差?
主要发现
- 该模型成功估计了不同站点间物种特异性生物量的变化,揭示了距道路距离的显著负向影响和海拔的正向影响,对多个物种具有统计显著性。
- 该框架识别出生物多样性高且物种生物量高的区域,可用于按保护价值对站点进行排序。
- 大多数物种的PCR错误概率接近于零,验证了本研究中所用实验室处理方案的可靠性。
- 模型检测到站点间及样本重复组内的噪声水平较低,表明数据采集和测序流程具有高度一致性。
- 自适应MCMC方法相比标准MCMC显著减少了计算时间,提升了大规模数据集的可扩展性。
- 通过重参数化,模型在复杂交叉效应结构中实现了更优的混合性和收敛性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。