Skip to main content
QUICK REVIEW

[論文レビュー] Detection of Risk Predictors of COVID-19 Mortality with Classifier Machine Learning Models Operated with Routine Laboratory Biomarkers

Mehmet Tahir Huyut, Andrei Velichko|arXiv (Cornell University)|Oct 22, 2022
COVID-19 diagnosis using AI被引用数 4
ひとこと要約

本研究では、2,597例のCOVID-19患者から得られた日常的血液検査バイオマーカーを用いて、ヒストグラムベースの勾配ブースティング(HGB)機械学習モデルを適用し、死亡リスクの予測要因を同定した。HGBモデルは、プロカルシトニンとDダイマー、ESR、D-ビリルビン、またはフェリチンの二重組み合わせを用いて、F1² > 0.98というほぼ完璧な分類性能を達成した。プロカルシトニン(F1² = 0.96)とフェリチン(F1² = 0.91)が単一特徴量として最も優れた結果を示し、死亡リスクが著しく上昇する臨床的臨界閾値として、プロカルシトニン 0.2–5.2 μg/L およびフェリチン 376.2–396.0 μg/L が特定された。

ABSTRACT

Early evaluation of patients who require special care and who have high death-expectancy in COVID-19, and the effective determination of relevant biomarkers on large sample-groups are important to reduce mortality. This study aimed to reveal the routine blood-value predictors of COVID-19 mortality and to determine the lethal-risk levels of these predictors during the disease process. The dataset of the study consists of 38 routine blood-values of 2597 patients who died (n = 233) and those who recovered (n = 2364) from COVID-19 in August-December, 2021. In this study, the histogram-based gradient-boosting (HGB) model was the most successful machine-learning classifier in detecting living and deceased COVID-19 patients (with squared F1 metrics F1^2 = 1). The most efficient binary combinations with procalcitonin were obtained with D-dimer, ESR, D-Bil and ferritin. The HGB model operated with these feature pairs correctly detected almost all of the patients who survived and those who died (precision > 0.98, recall > 0.98, F1^2 > 0.98). Furthermore, in the HGB model operated with a single feature, the most efficient features were procalcitonin (F1^2 = 0.96) and ferritin (F1^2 = 0.91). In addition, according to the two-threshold approach, ferritin values between 376.2 mkg/L and 396.0 mkg/L (F1^2 = 0.91) and pro-calcitonin values between 0.2 mkg/L and 5.2 mkg/L (F1^2 = 0.95) were found to be fatal risk levels for COVID-19. Considering all the results, we suggest that many features combined with these features, especially procalcitonin and ferritin, operated with the HGB model, can be used to achieve very successful results in the classification of those who live, and those who die from COVID-19. Moreover, we strongly recommend that clinicians consider the critical levels we have found for procalcitonin and ferritin properties, to reduce the lethality of the COVID-19 disease.

研究の動機と目的

  • 機械学習を用いて、COVID-19死亡リスクを予測可能な日常的血液検査バイオマーカーを同定すること。
  • 入院中のCOVID-19患者の生存・死亡を分類するための最適な特徴量の組み合わせを特定すること。
  • 疾病進行中に致命的リスクを示す重要なバイオマーカーの閾値を確立すること。
  • 標準的な血液検査データを用いて、生存者と非生存者を区別する複数の分類器モデルの性能を評価すること。
  • 早期リスク評価を支援し、死亡率を低減するための臨床的実用可能なバイオマーカーの閾値を提供すること。

提案手法

  • 不均衡データに対して高い性能を示すため、主にヒストグラムベースの勾配ブースティング(HGB)を分類器として採用した。
  • 2021年8月から12月にかけて収集された2,597名の患者から得られた38種類の日常的血液バイオマーカーを用い、プロカルシトニン、フェリチン、Dダイマー、ESR、ビリルビンを含む。
  • 分類精度を評価するため、F1²スコア、適合率、再現率、受信者動作特性曲線下の面積(AUC)を用いてモデルの性能を評価した。
  • 臨床的関連性のあるバイオマーカー範囲を同定するため、二重閾値アプローチを採用した。
  • 単一特徴量モデルと二重組み合わせモデルを比較し、最適な予測ペアおよび最も優れた個別特徴量を特定した。
  • 10分割交差検証を用いて結果を検証し、データセット全体にわたるモデルの汎用性についての信頼性を報告した。

実験結果

リサーチクエスチョン

  • RQ1機械学習モデルを用いた場合、どの日常的血液バイオマーカーがCOVID-19死亡リスクに対して最も予測的であるか?
  • RQ2生存と死亡を分類する際の分類精度を最大化する2つのバイオマーカーの最適な組み合わせは何か?
  • RQ3HGBモデルを用いた場合、死亡リスクを予測するうえで最も高いF1²スコアを示す単一バイオマーカーは何か?
  • RQ4死亡リスクが著しく上昇するプロカルシトニンおよびフェリチンの臨界閾値範囲は何か?
  • RQ5標準的な血液検査データを用いた機械学習モデルは、疾病進行の初期段階で高リスクと低リスクの患者を信頼性高く区別できるか?

主な発見

  • HGBモデルは、プロカルシトニンとDダイマー、ESR、D-ビリルビン、またはフェリチンの二重組み合わせを用いることで、F1² = 1.0という最高の性能を達成した。
  • プロカルシトニン単体ではF1²スコアが0.96を達成し、死亡リスク予測において最も効果的な単一バイオマーカーとなった。
  • フェリチン単体でもF1²スコアが0.91を達成し、致命的結果の予測能が非常に高いことが示された。
  • 二重閾値アプローチにより、プロカルシトニン濃度が0.2–5.2 μg/Lの範囲が、F1² = 0.95で高いリスク範囲と特定された。
  • フェリチン濃度が376.2–396.0 μg/Lの範囲は、F1² = 0.91で致命的リスクの臨界閾値として特定された。
  • 最良の特徴量ペアを用いたHGBモデルは、適合率 > 0.98、再現率 > 0.98、F1² > 0.98を達成し、生存者と非生存者のほぼ完璧な分類が可能となった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。