Skip to main content
QUICK REVIEW

[論文レビュー] Interpretable Machine Learning Model for Early Prediction of Mortality in Elderly Patients with Multiple Organ Dysfunction Syndrome (MODS): a Multicenter Retrospective Study and Cross Validation

Xiaoli Liu, Pan Hu|arXiv (Cornell University)|Jan 28, 2020
Machine Learning in Healthcare参考文献 7被引用数 5
ひとこと要約

本研究では、複数のICUデータベース(MIMIC-III、eICU-CRD、PLAGH-S)を用いた多施設のICUデータを活用し、高齢の多臓器不全症候群(MODS)患者における入院死亡の早期予測を目的とした解釈可能なXGBoostベースの機械学習モデルを開発した。3つの独立したデータセットにおいて、AUCがそれぞれ0.858、0.849、0.838を達成し、従来のスコアリングシステムやベースラインモデルを上回った。また、SHAPによる解釈可能性を提供し、特徴量の重要度を明らかにした。

ABSTRACT

Background: Elderly patients with MODS have high risk of death and poor prognosis. The performance of current scoring systems assessing the severity of MODS and its mortality remains unsatisfactory. This study aims to develop an interpretable and generalizable model for early mortality prediction in elderly patients with MODS. Methods: The MIMIC-III, eICU-CRD and PLAGH-S databases were employed for model generation and evaluation. We used the eXtreme Gradient Boosting model with the SHapley Additive exPlanations method to conduct early and interpretable predictions of patients' hospital outcome. Three types of data source combinations and five typical evaluation indexes were adopted to develop a generalizable model. Findings: The interpretable model, with optimal performance developed by using MIMIC-III and eICU-CRD datasets, was separately validated in MIMIC-III, eICU-CRD and PLAGH-S datasets (no overlapping with training set). The performances of the model in predicting hospital mortality as validated by the three datasets were: AUC of 0.858, sensitivity of 0.834 and specificity of 0.705; AUC of 0.849, sensitivity of 0.763 and specificity of 0.784; and AUC of 0.838, sensitivity of 0.882 and specificity of 0.691, respectively. Comparisons of AUC between this model and baseline models with MIMIC-III dataset validation showed superior performances of this model; In addition, comparisons in AUC between this model and commonly used clinical scores showed significantly better performance of this model. Interpretation: The interpretable machine learning model developed in this study using fused datasets with large sample sizes was robust and generalizable. This model outperformed the baseline models and several clinical scores for early prediction of mortality in elderly ICU patients. The interpretative nature of this model provided clinicians with the ranking of mortality risk features.

研究の動機と目的

  • 高齢の多臓器不全症候群(MODS)患者における高い死亡リスクと不良な予後を改善すること。
  • 従来の臨床スコアリングシステムがMODS関連死亡を予測する際の限界を克服すること。
  • 高齢ICU患者における早期死亡予測のための解釈可能で汎用性のある機械学習モデルを開発すること。
  • 複数の多施設ICUデータベースからの統合データを活用し、モデルの強固さと外部妥当性を向上させること。

提案手法

  • 大規模なICUデータセットを対象に予測モデリングにeXtreme Gradient Boosting(XGBoost)アルゴリズムを適用した。
  • モデルの予測を解釈し、特徴量の重要度を順位付けするためにSHapley Additive exPlanations(SHAP)を適用した。
  • MIMIC-III、eICU-CRD、PLAGH-Sの3つの異なるICUデータベースのデータを統合し、モデルの学習と交差検証に用いた。
  • AUC、感度、特異度を含む5つの評価指標を用いて、モデルのパフォーマンスを評価した。
  • 汎用性を確保するため、訓練データとテストデータの間に重複する患者が存在しないように、データセット間での検証を実施した。
  • MIMIC-IIIおよびeICU-CRDデータセットを用いてモデルのパフォーマンスを最適化し、その後、3つのデータセットすべてに対して独立した検証を実施した。

実験結果

リサーチクエスチョン

  • RQ1従来の臨床スコアリングシステムと比較して、解釈可能な機械学習モデルは、高齢のMODS患者における入院死亡の早期予測を改善できるか?
  • RQ2複数の独立した多施設ICUデータベースで検証された場合、モデルの汎用性はどの程度か?
  • RQ3SHAPによる解釈可能性によって明らかにされた臨床的特徴量の中で、死亡予測に最も寄与しているのはどれか?
  • RQ4複数のデータソースからの統合は、MODS死亡予測におけるモデルのパフォーマンスと強固さを向上させるか?

主な発見

  • MIMIC-IIIデータセットでは、AUCが0.858に達し、感度は0.834、特異度は0.705を示した。
  • eICU-CRDデータセットでは、AUCが0.849に達し、感度は0.763、特異度は0.784を示した。
  • PLAGH-Sデータセットでは、AUCが0.838に達し、感度は0.882、特異度は0.691を示した。
  • すべての検証データセットにおいて、AUCの観点で、ベースラインモデルおよび一般的に使用される臨床スコアを有意に上回った。
  • SHAP解析により、死亡リスク要因の解釈可能な順位付けが可能となり、臨床的信頼性と実用性が向上した。
  • 重複のない多様なICUデータセットにおいて、モデルが強く汎用可能であることが確認され、強固さと外部妥当性が裏付けられた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。