Skip to main content
QUICK REVIEW

[論文レビュー] Data-driven Identification of Number of Unreported Cases for COVID-19: Bounds and Limitations

Ajitesh Srivastava, Viktor K. Prasanna|arXiv (Cornell University)|Jun 3, 2020
Anomaly Detection Techniques and Applications参考文献 9被引用数 4
ひとこと要約

本稿では、安定した社会的距離の維持段階における疫学的データを用いて、報告されていないCOVID-19症例数の上限を推定するデータ駆動型手法、Fixed Infection Rate Learningを提案する。報告確率が特定の時間窓でのみ信頼性を持って学習可能であることを証明することで、実際の症例数と報告された症例数の比の上限を制限する。ニューヨークでは35倍以下、イリノイ州では40倍以下、マサチューセッツ州では38倍以下、ニュージャージー州では29倍以下であり、高い信頼性を伴う。

ABSTRACT

Accurate forecasts for COVID-19 are necessary for better preparedness and resource management. Specifically, deciding the response over months or several months requires accurate long-term forecasts which is particularly challenging as the model errors accumulate with time. A critical factor that can hinder accurate long-term forecasts, is the number of unreported/asymptomatic cases. While there have been early serology tests to estimate this number, more tests need to be conducted for more reliable results. To identify the number of unreported/asymptomatic cases, we take an epidemiology data-driven approach. We show that we can identify lower bounds on this ratio or upper bound on actual cases as a factor of reported cases. To do so, we propose an extension of our prior heterogeneous infection rate model, incorporating unreported/asymptomatic cases. We prove that the number of unreported cases can be reliably estimated only from a certain time period of the epidemic data. In doing so, we construct an algorithm called Fixed Infection Rate method, which identifies a reliable bound on the learned ratio. We also propose two heuristics to learn this ratio and show their effectiveness on simulated data. We use our approaches to identify the upper bounds on the ratio of actual to reported cases for New York City and several US states. Our results demonstrate with high confidence that the actual number of cases cannot be more than 35 times in New York, 40 times in Illinois, 38 times in Massachusetts and 29 times in New Jersey, than the reported cases.

研究の動機と目的

  • 報告されていないまたは無症状の症例による長期予測の信頼性の欠如という課題に対処すること。
  • 報告確率を最小限の誤差で推定可能な、疫学的データにおける信頼性の高い時間窓を同定すること。
  • 未知の隔離効果が存在する中でも、実際の症例数と報告された症例数の比の保証された上限を提供する手法を開発すること。
  • 米国各州の実世界データを用いて、本手法の性能を評価し、ヒューリスティックな代替手法と比較すること。

提案手法

  • 報告された症例数と実際の症例数の比を表すパラメータを含む、先行研究の非一様感染率モデルを拡張する。
  • 初期の不安定性が解消された後で、終末段階の動的変化が現れる前の特定の時間間隔を同定し、報告確率を信頼性を持って推定可能な時期とする。
  • Fixed Infection Rate Learningアルゴリズムを提案する。この手法は、この時間窓におけるデータの安定性を活用し、報告確率の下限を計算することで、総症例数の上限を導出する。
  • 理論的保証のない2つのヒューリスティック、Non-linear Incremental LearningおよびNon-linear Curve Fittingを導入し、境界を推定する。
  • 報告確率と人口の隔離(γ̄ = (1−ρ)γ)の組み合わせ効果を代理指標として用いる。ここでρは完全に隔離されている人口の割合を表す。
  • ニューヨーク市および米国複数州の実データに本手法を適用し、統計的検定と信頼区間を用いて結果の妥当性を検証する。

実験結果

リサーチクエスチョン

  • RQ1社会的距離の維持段階における疫学的データから、報告されていないCOVID-19症例数の信頼性の高い上限を導出できるか?
  • RQ2モデルの不安定性が顕著な初期段階と、終末段階での動的変化が現れる時期の間で、報告確率を最も正確に推定可能な時間窓は何か?
  • RQ3隔離されている人々(感染伝播も報告も行わない者)の存在が、報告確率の学習可能性および総症例数の上限に与える影響は何か?
  • RQ4提案手法は、ヒューリスティックな代替手法に比べて、実際の症例数の上限を推定する上でどの程度優れているか?
  • RQ5本手法は、他の疫学的モデルに一般化可能であり、伝染の傾向が多様な地域に応用可能か?

主な発見

  • ニューヨーク州におけるCOVID-19の実際の症例数は、報告された症例数の35倍を超えることはない(高い信頼性を伴う)。
  • イリノイ州では、総症例数の上限が報告された症例数の40倍以下である。
  • マサチューセッツ州では、上限が報告された症例数の38倍以下である。
  • ニュージャージー州では、上限が報告された症例数の29倍以下である。
  • 4州全員が統計的検定(Test1およびTest2)に合格し、推定された上限の信頼性が確認された。
  • ロサンゼルスでは、95%信頼区間が妥当範囲外に位置していたため、本手法の性能は一貫しなかった。これは、ここではまだ境界を信頼性を持って推定するには時期が早い可能性を示唆している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。