Skip to main content
QUICK REVIEW

[论文解读] Interpretable Machine Learning Model for Early Prediction of Mortality in Elderly Patients with Multiple Organ Dysfunction Syndrome (MODS): a Multicenter Retrospective Study and Cross Validation

Xiaoli Liu, Pan Hu|arXiv (Cornell University)|Jan 28, 2020
Machine Learning in Healthcare参考文献 7被引用 5
一句话总结

本研究基于多中心ICU数据库(MIMIC-III、eICU-CRD、PLAGH-S)开发了一种可解释的基于XGBoost的机器学习模型,用于早期预测老年多器官功能障碍综合征(MODS)患者的住院死亡率。该模型在三个独立数据集中的AUC分别为0.858、0.849和0.838,优于传统的评分系统和基线模型,同时通过SHAP解释提供了特征重要性排序。

ABSTRACT

Background: Elderly patients with MODS have high risk of death and poor prognosis. The performance of current scoring systems assessing the severity of MODS and its mortality remains unsatisfactory. This study aims to develop an interpretable and generalizable model for early mortality prediction in elderly patients with MODS. Methods: The MIMIC-III, eICU-CRD and PLAGH-S databases were employed for model generation and evaluation. We used the eXtreme Gradient Boosting model with the SHapley Additive exPlanations method to conduct early and interpretable predictions of patients' hospital outcome. Three types of data source combinations and five typical evaluation indexes were adopted to develop a generalizable model. Findings: The interpretable model, with optimal performance developed by using MIMIC-III and eICU-CRD datasets, was separately validated in MIMIC-III, eICU-CRD and PLAGH-S datasets (no overlapping with training set). The performances of the model in predicting hospital mortality as validated by the three datasets were: AUC of 0.858, sensitivity of 0.834 and specificity of 0.705; AUC of 0.849, sensitivity of 0.763 and specificity of 0.784; and AUC of 0.838, sensitivity of 0.882 and specificity of 0.691, respectively. Comparisons of AUC between this model and baseline models with MIMIC-III dataset validation showed superior performances of this model; In addition, comparisons in AUC between this model and commonly used clinical scores showed significantly better performance of this model. Interpretation: The interpretable machine learning model developed in this study using fused datasets with large sample sizes was robust and generalizable. This model outperformed the baseline models and several clinical scores for early prediction of mortality in elderly ICU patients. The interpretative nature of this model provided clinicians with the ranking of mortality risk features.

研究动机与目标

  • 为解决老年多器官功能障碍综合征(MODS)患者高死亡率和预后不良的问题。
  • 克服现有临床评分系统在预测MODS相关死亡率方面的局限性。
  • 开发一种可解释、可泛化的机器学习模型,用于老年ICU患者早期死亡率预测。
  • 整合来自多个多中心ICU数据库的数据,以增强模型的稳健性和外部有效性。

提出的方法

  • 采用极端梯度提升(XGBoost)算法,在大规模ICU数据集上进行预测建模。
  • 应用SHapley加性解释(SHAP)以解释模型预测结果并排序特征重要性。
  • 整合来自三个不同ICU数据库——MIMIC-III、eICU-CRD和PLAGH-S——的数据用于模型训练和交叉验证。
  • 采用五项评估指标,包括AUC、敏感性和特异性,以评估模型性能。
  • 通过跨数据集验证确保模型的泛化能力,训练集与测试集之间无重叠患者。
  • 使用MIMIC-III和eICU-CRD数据集优化模型性能,随后在所有三个数据集上独立验证。

实验结果

研究问题

  • RQ1与现有临床评分系统相比,可解释的机器学习模型是否能提升对老年MODS患者住院死亡率的早期预测能力?
  • RQ2当在多个独立的多中心ICU数据库上验证时,该模型的泛化能力如何?
  • RQ3通过SHAP可解释性分析,哪些临床特征对死亡率预测的贡献最为显著?
  • RQ4结合多个来源的数据是否能提升模型在预测MODS死亡率方面的性能和稳健性?

主要发现

  • 在MIMIC-III数据集中,模型AUC为0.858,敏感性为0.834,特异性为0.705。
  • 在eICU-CRD数据集中,模型AUC为0.849,敏感性为0.763,特异性为0.784。
  • 在PLAGH-S数据集中,模型AUC为0.838,敏感性为0.882,特异性为0.691。
  • 在所有验证数据集中,该模型在AUC方面显著优于基线模型和常用临床评分。
  • SHAP分析提供了可解释的死亡风险特征排序,增强了临床信任度和实用性。
  • 该模型在多样化、无重叠的ICU数据集中表现出强大的泛化能力,证实了其稳健性和外部有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。