Skip to main content
QUICK REVIEW

[论文解读] Bayesian adjustment for preferential testing in estimating the COVID-19 infection fatality rate: Theory and methods

Harlan Campbell, Valpine Pd|arXiv (Cornell University)|May 18, 2020
COVID-19 epidemiological studies参考文献 48被引用 13
一句话总结

本文提出了一种贝叶斯分层模型,用于在估计COVID-19感染致死率(IFR)时校正偏好性检测——即感染个体更可能被检测——的影响。通过整合来自多个具有不同程度检测偏差的数据源,并利用关于偏差异质性的先验假设,该模型实现了可靠的IFR估计;将其应用于欧洲数据后,估计总体IFR为0.47%(95%可信区间:0.34%–0.63%)。

ABSTRACT

A key challenge in estimating the infection fatality rate (IFR) is determining the total number of cases. The total number of cases is not known because not everyone is tested but also, more importantly, because tested individuals are not representative of the population at large. We refer to the phenomenon whereby infected individuals are more likely to be tested than non-infected individuals, as ''preferential testing.'' An open question is whether or not it is possible to reliably estimate the IFR without any specific knowledge about the degree to which the data are biased by preferential testing. In this paper we take a partial identifiability approach, formulating clearly where deliberate prior assumptions can be made and presenting a Bayesian model, which pools information from different samples. Results of a simulation study suggest that when certain populations with representative testing (i.e., or for which the degree of preferential testing is known) are included in the analysis, identifiability and reliable estimates are attainable. When only limited knowledge is available about the magnitude of preferential testing, reliable estimation of the IFR may still be possible so long as there is sufficient ''heterogeneity of bias'' across samples. When the model is fit to European data obtained from seroprevalence studies and national official COVID-19 statistics, we estimate the overall COVID-19 IFR for Europe to be 0.47%, 95% C.I. = [0.34%, 0.63%].

研究动机与目标

  • 解决在估计COVID-19感染致死率(IFR)时,因感染个体更可能被检测而导致的偏好性检测问题。
  • 确定在缺乏测试偏差完全知识的情况下,是否仍可实现可靠的IFR估计,基于部分可识别性框架。
  • 开发一种贝叶斯模型,汇集来自具有不同程度偏好性检测的多个样本的数据,以提高估计精度。
  • 评估在缺乏偏差完全知识的情况下,识别IFR并获得可靠IFR估计的条件。
  • 将该模型应用于现实世界中的欧洲血清流行率和官方统计数据,为该地区生成稳健的IFR估计。

提出的方法

  • 作者采用一种贝叶斯分层模型,考虑不同人群间检测概率的差异,将偏好性检测视为一种选择偏差来源。
  • 该模型在测试偏差程度上引入先验分布,使得在缺乏完整偏差知识时仍能实现部分可识别性。
  • 它汇集来自多个样本的数据,包括一些具有已知或代表性检测行为的样本,以提高估计的稳定性并减少偏差。
  • 该方法依赖于样本间偏差的异质性,即使单个样本的偏差不确定,也能实现IFR的识别。
  • 该模型基于观察到的病例数和死亡数构建似然函数,并通过来自辅助数据的估计检测概率进行调整。
  • 通过马尔可夫链蒙特卡洛(MCMC)抽样进行后验推断,以估计IFR及其可信区间。

实验结果

研究问题

  • RQ1当检测存在偏好性且偏差程度未知时,是否可以可靠地估计COVID-19的感染致死率(IFR)?
  • RQ2在样本间测试偏差存在异质性的条件下,IFR的可识别性在什么情况下能够实现?
  • RQ3将具有已知或代表性检测行为的人群数据纳入分析,如何提升IFR估计的准确性?
  • RQ4关于偏好性检测的先验假设对最终IFR估计有何影响?
  • RQ5当将具有部分可识别性的贝叶斯模型应用于现实世界中的欧洲数据时,其IFR估计在多大程度上具有稳健性?

主要发现

  • 该贝叶斯模型成功估计出欧洲总体的COVID-19感染致死率(IFR)为0.47%,95%可信区间为[0.34%,0.63%]。
  • 当至少部分样本具有已知或代表性检测行为时,即使其他样本的偏差未知,仍可实现可靠的IFR估计。
  • 若样本间偏好性检测程度存在足够异质性,即使缺乏完整的偏差信息,也能实现IFR的识别。
  • 模拟研究证实,该模型在现实偏差和数据可用性场景下保持了良好的频率学性质。
  • 纳入代表性检测数据显著提高了IFR估计的精度和可靠性。
  • 当先验分布通过辅助数据适当地加以信息化时,该模型对偏差参数的不确定性表现出稳健性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。