[论文解读] Testing Case Number of Coronavirus Disease 2019 in China with Newcomb-Benford Law
本研究使用Newcomb-Benford定律(NBL),一种用于检测数据欺诈的统计方法,检验中国报告的COVID-19累计病例数的完整性。结果得到92.8%的p值,表明与NBL高度一致,且在疫情早期阶段的报告病例数中无统计证据显示存在数据操纵行为。
The coronavirus disease 2019 bursted out about two months ago in Wuhan has caused the death of more than a thousand people. China is fighting hard against the epidemics with the helps from all over the world. On the other hand, there appear to be doubts on the reported case number. In this article, we propose a test of the reported case number of coronavirus disease 2019 in China with Newcomb-Benford law. We find a $p$-value of $92.8\%$ in favour that the cumulative case numbers abide by the Newcomb-Benford law. Even though the reported case number can be lower than the real number of affected people due to various reasons, this test does not seem to indicate the detection of frauds.
研究动机与目标
- 评估中国在疫情早期阶段报告的COVID-19累计病例数的统计完整性。
- 确定病例数据中首位有效数字的分布是否符合Newcomb-Benford定律(NBL)。
- 评估观察到的数据模式是否暗示可能存在数据伪造或操纵行为。
- 探讨NBL作为检测流行病统计数据异常的工具的适用性。
提出的方法
- 本研究将Newcomb-Benford定律(NBL)应用于中国31个省级行政区累计确诊病例数的首位有效数字。
- 使用卡方检验统计量 $ V = \sum_{d=1}^{9} \frac{(N(d) - N_{\text{tot}} P_{\text{NB}}(d))^2}{N_{\text{tot}} P_{\text{NB}}(d)} $,将观察到的数字频率与NBL的期望值进行比较。
- 原假设 $ H_0 $ 假设数据符合NBL,意味着无欺诈行为;备择假设则包含通过混合模型 $ \Pi = (1-\tau)\Psi + \tau\Phi $ 潜在欺诈的可能性。
- 该测试基于2020年1月15日至2月10日期间从维基百科获取的628个公开数据点进行。
- 计算得出92.8%的p值,以评估观察到的分布与NBL期望分布之间的拟合优度。
实验结果
研究问题
- RQ1中国报告的累计COVID-19病例数中,首位有效数字的分布是否符合Newcomb-Benford定律?
- RQ2在疫情早期阶段,报告病例数中是否存在统计证据表明存在数据操纵或欺诈行为?
- RQ3Newcomb-Benford定律能否可靠地应用于流行病病例数据以检测异常?
- RQ4当将NBL测试应用于公共卫生情境下的早期指数增长数据时,其稳健性如何?
主要发现
- 卡方检验统计量得出 $ V = 3.10 $,对应p值为92.8%。
- 较高的p值表明与Newcomb-Benford定律高度一致,支持原假设,即数据遵循预期的数字分布。
- 在研究期间,中国报告的COVID-19累计病例数中,无统计证据表明存在数据欺诈或操纵行为。
- 该测试无法排除由于医疗资源有限导致的潜在漏报或系统性偏差,因为NBL无法检测系统性漏报。
- 结果表明,NBL可作为初步审查流行病数据完整性的有用工具,但并非能检测所有形式数据偏差的决定性方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。