[论文解读] Data-driven Identification of Number of Unreported Cases for COVID-19: Bounds and Limitations
本文提出一种数据驱动方法——固定感染率学习,利用稳定社交距离阶段的流行病数据,估算未报告的COVID-19病例数的上界。通过证明报告概率仅能在特定时间窗口内被可靠学习,该方法界定了实际病例与报告病例之比的上界——在纽约州,该比例不能超过35倍;在伊利诺伊州,不能超过40倍;在马萨诸塞州,不能超过38倍;在新泽西州,不能超过29倍,且置信度较高。
Accurate forecasts for COVID-19 are necessary for better preparedness and resource management. Specifically, deciding the response over months or several months requires accurate long-term forecasts which is particularly challenging as the model errors accumulate with time. A critical factor that can hinder accurate long-term forecasts, is the number of unreported/asymptomatic cases. While there have been early serology tests to estimate this number, more tests need to be conducted for more reliable results. To identify the number of unreported/asymptomatic cases, we take an epidemiology data-driven approach. We show that we can identify lower bounds on this ratio or upper bound on actual cases as a factor of reported cases. To do so, we propose an extension of our prior heterogeneous infection rate model, incorporating unreported/asymptomatic cases. We prove that the number of unreported cases can be reliably estimated only from a certain time period of the epidemic data. In doing so, we construct an algorithm called Fixed Infection Rate method, which identifies a reliable bound on the learned ratio. We also propose two heuristics to learn this ratio and show their effectiveness on simulated data. We use our approaches to identify the upper bounds on the ratio of actual to reported cases for New York City and several US states. Our results demonstrate with high confidence that the actual number of cases cannot be more than 35 times in New York, 40 times in Illinois, 38 times in Massachusetts and 29 times in New Jersey, than the reported cases.
研究动机与目标
- 解决因未报告或无症状病例导致的长期流行病预测不可靠的问题。
- 识别流行病数据中一个可靠的时段窗口,在此期间报告概率可被最小误差地估计。
- 开发一种方法,即使在隔离效应未知的情况下,也能提供实际病例与报告病例之比的保证上界。
- 在真实世界美国各州数据上评估该方法的性能,并与启发式方法进行比较。
提出的方法
- 将先前的异质感染率模型扩展,引入一个表示报告病例与实际病例之比的参数。
- 识别流行病进程中一个特定时间段——初始不稳定期过后、晚期动态出现前——在此期间报告概率可被可靠估计。
- 提出固定感染率学习算法,利用该时间段内的数据稳定性,计算报告概率的下界,从而得到总病例数的上界。
- 引入两种启发式方法——非线性增量学习与非线性曲线拟合——用于估计上界,但不提供理论保证。
- 使用报告概率与人群隔离效应(γ̄ = (1−ρ)γ)的综合影响作为代理变量,其中ρ为完全隔离人群所占比例。
- 将该方法应用于纽约市及多个美国州的真实数据,通过统计检验和置信区间验证结果。
实验结果
研究问题
- RQ1能否从社交距离阶段的流行病数据中,推导出未报告的COVID-19病例数的可靠上界?
- RQ2在流行病进程中,哪个时间窗口最有利于准确估计报告概率?(考虑到初期模型不稳定,后期动态发生变化。)
- RQ3隔离人群(不传播病毒也不报告)的存在如何影响报告概率的学习性,以及对总病例数上界估计的影响?
- RQ4所提出的方法在估计实际病例数上界方面,相较于启发式替代方法,优势有多大?
- RQ5该方法能否推广至其他流行病学模型,并应用于具有异质传播趋势的不同地区?
主要发现
- 在纽约州,实际感染的COVID-19病例数不可能超过报告病例数的35倍,且置信度较高。
- 在伊利诺伊州,总病例数的上界为报告病例数的40倍。
- 在马萨诸塞州,总病例数的上界为报告病例数的38倍。
- 在新泽西州,总病例数的上界为报告病例数的29倍。
- 上述四个州均通过了两项统计检验(Test1 和 Test2),证实了所估计上界的可靠性。
- 在洛杉矶,该方法表现不稳定,95%置信区间超出可行范围,表明该地可能尚且过早,无法可靠估计上界。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。