Skip to main content
QUICK REVIEW

[论文解读] Detection of Risk Predictors of COVID-19 Mortality with Classifier Machine Learning Models Operated with Routine Laboratory Biomarkers

Mehmet Tahir Huyut, Andrei Velichko|arXiv (Cornell University)|Oct 22, 2022
COVID-19 diagnosis using AI被引用 4
一句话总结

本研究利用基于直方图的梯度提升(HGB)机器学习模型,对2,597名COVID-19患者的常规实验室生物标志物进行分析,以识别死亡预测因子。HGB模型在使用降钙素原与D-二聚体、血沉(ESR)、直接胆红素(D-Bil)或铁蛋白的二元组合时,实现了近乎完美的分类效果(F1² > 0.98),其中降钙素原(F1² = 0.96)和铁蛋白(F1² = 0.91)作为单一特征表现最佳,同时确定了关键风险阈值:降钙素原0.2–5.2 μg/L 和铁蛋白376.2–396.0 μg/L。

ABSTRACT

Early evaluation of patients who require special care and who have high death-expectancy in COVID-19, and the effective determination of relevant biomarkers on large sample-groups are important to reduce mortality. This study aimed to reveal the routine blood-value predictors of COVID-19 mortality and to determine the lethal-risk levels of these predictors during the disease process. The dataset of the study consists of 38 routine blood-values of 2597 patients who died (n = 233) and those who recovered (n = 2364) from COVID-19 in August-December, 2021. In this study, the histogram-based gradient-boosting (HGB) model was the most successful machine-learning classifier in detecting living and deceased COVID-19 patients (with squared F1 metrics F1^2 = 1). The most efficient binary combinations with procalcitonin were obtained with D-dimer, ESR, D-Bil and ferritin. The HGB model operated with these feature pairs correctly detected almost all of the patients who survived and those who died (precision > 0.98, recall > 0.98, F1^2 > 0.98). Furthermore, in the HGB model operated with a single feature, the most efficient features were procalcitonin (F1^2 = 0.96) and ferritin (F1^2 = 0.91). In addition, according to the two-threshold approach, ferritin values between 376.2 mkg/L and 396.0 mkg/L (F1^2 = 0.91) and pro-calcitonin values between 0.2 mkg/L and 5.2 mkg/L (F1^2 = 0.95) were found to be fatal risk levels for COVID-19. Considering all the results, we suggest that many features combined with these features, especially procalcitonin and ferritin, operated with the HGB model, can be used to achieve very successful results in the classification of those who live, and those who die from COVID-19. Moreover, we strongly recommend that clinicians consider the critical levels we have found for procalcitonin and ferritin properties, to reduce the lethality of the COVID-19 disease.

研究动机与目标

  • 利用机器学习识别常规实验室生物标志物对COVID-19死亡的预测能力。
  • 确定用于区分住院COVID-19患者生存与死亡的最有效特征组合。
  • 确定在疾病进展过程中提示高致死风险的关键生物标志物临界阈值水平。
  • 评估多种分类器模型在使用标准实验室数据区分幸存者与非幸存者方面的性能。
  • 提供临床可操作的生物标志物阈值,以支持早期风险分层并降低死亡率。

提出的方法

  • 由于HGB在处理不平衡数据方面表现优异,故将其作为主要机器学习分类器。
  • 使用2021年8月至12月期间从2,597名患者中收集的38种常规血液生物标志物,包括降钙素原、铁蛋白、D-二聚体、血沉(ESR)和胆红素。
  • 采用F1²分数、精确率、召回率以及受试者工作特征曲线下面积(AUC)评估模型性能,以衡量分类准确性。
  • 采用双阈值方法识别与高死亡风险相关的临床相关生物标志物范围。
  • 比较单特征与二元组合模型,以确定最优预测组合对及表现最佳的单一特征。
  • 通过10折交叉验证验证结果,并报告模型在数据集上泛化能力的置信度。

实验结果

研究问题

  • RQ1在机器学习模型中,哪些常规实验室生物标志物对COVID-19死亡最具预测性?
  • RQ2哪两种生物标志物的组合能最大化区分COVID-19患者生存与死亡的分类准确率?
  • RQ3在HGB模型中,哪种单一生物标志物在预测死亡时获得最高的F1²分数?
  • RQ4降钙素原与铁蛋白的哪些临界阈值范围与死亡风险显著增加相关?
  • RQ5使用标准实验室检测的机器学习模型能否可靠地区分疾病早期的高危与低危患者?

主要发现

  • 当使用降钙素原与D-二聚体、ESR、D-Bil或铁蛋白的二元组合时,HGB模型的F1²达到最高值1.0。
  • 单独使用降钙素原时,F1²得分为0.96,表明其作为单一生物标志物在死亡预测中最为有效。
  • 单独使用铁蛋白时,F1²得分为0.91,表明其对不良结局具有强大的预测能力。
  • 双阈值方法识别出降钙素原水平在0.2至5.2 μg/L之间为高风险范围,F1²为0.95。
  • 铁蛋白水平在376.2至396.0 μg/L之间被确定为关键致死风险阈值,F1²为0.91。
  • 使用最佳特征对的HGB模型实现了精确率 > 0.98、召回率 > 0.98 以及 F1² > 0.98,表明对幸存者与非幸存者的分类近乎完美。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。