Skip to main content
QUICK REVIEW

[논문 리뷰] Interpretable Machine Learning Model for Early Prediction of Mortality in Elderly Patients with Multiple Organ Dysfunction Syndrome (MODS): a Multicenter Retrospective Study and Cross Validation

Xiaoli Liu, Pan Hu|arXiv (Cornell University)|2020. 01. 28.
Machine Learning in Healthcare참고 문헌 7인용 수 5
한 줄 요약

이 연구는 다기관 중환자실 데이터베이스(MIMIC-III, eICU-CRD, PLAGH-S)를 활용하여 노인 환자에서 다중 기관 기능 저하 증후군(MODS)의 조기 병원 사망을 예측하기 위한 해석 가능한 XGBoost 기반 기계학습 모델을 개발한다. 모델은 세 개의 독립된 데이터셋에서 AUC 0.858, 0.849, 0.838을 기록하며 기존의 임상 점수 체계와 기준 모델을 능가한다. 동시에 SHAP 해석을 통해 특성 중요도를 제공한다.

ABSTRACT

Background: Elderly patients with MODS have high risk of death and poor prognosis. The performance of current scoring systems assessing the severity of MODS and its mortality remains unsatisfactory. This study aims to develop an interpretable and generalizable model for early mortality prediction in elderly patients with MODS. Methods: The MIMIC-III, eICU-CRD and PLAGH-S databases were employed for model generation and evaluation. We used the eXtreme Gradient Boosting model with the SHapley Additive exPlanations method to conduct early and interpretable predictions of patients' hospital outcome. Three types of data source combinations and five typical evaluation indexes were adopted to develop a generalizable model. Findings: The interpretable model, with optimal performance developed by using MIMIC-III and eICU-CRD datasets, was separately validated in MIMIC-III, eICU-CRD and PLAGH-S datasets (no overlapping with training set). The performances of the model in predicting hospital mortality as validated by the three datasets were: AUC of 0.858, sensitivity of 0.834 and specificity of 0.705; AUC of 0.849, sensitivity of 0.763 and specificity of 0.784; and AUC of 0.838, sensitivity of 0.882 and specificity of 0.691, respectively. Comparisons of AUC between this model and baseline models with MIMIC-III dataset validation showed superior performances of this model; In addition, comparisons in AUC between this model and commonly used clinical scores showed significantly better performance of this model. Interpretation: The interpretable machine learning model developed in this study using fused datasets with large sample sizes was robust and generalizable. This model outperformed the baseline models and several clinical scores for early prediction of mortality in elderly ICU patients. The interpretative nature of this model provided clinicians with the ranking of mortality risk features.

연구 동기 및 목표

  • 노인 환자에서 다중 기관 기능 저하 증후군(MODS)의 높은 사망 위험도와 악화된 예후를 해결하기 위해.
  • 기존 임상 점수 체계가 MODS 관련 사망을 예측하는 데에 한계를 보이고 있는 문제를 해결하기 위해.
  • 노인 중환자실 환자에서 조기 사망 예측을 위한 해석 가능하고 일반화 가능한 기계학습 모델을 개발하기 위해.
  • 다양한 다기관 중환자실 데이터베이스의 데이터를 통합하여 모델의 강건성과 외부 타당성을 향상시키기 위해.

제안 방법

  • 대규모 중환자실 데이터셋에서 예측 모델링을 위해 초고속 기반 기계학습 알고리즘인 XGBoost를 사용하였다.
  • 모델 예측의 해석과 특성 중요도 순위 매기기를 위해 SHapley Additive exPlanations(SHAP)를 적용하였다.
  • 모델 훈련 및 교차 검증을 위해 MIMIC-III, eICU-CRD, PLAGH-S의 세 가지 서로 다른 중환자실 데이터베이스의 데이터를 통합하였다.
  • 모델 성능 평가를 위해 AUC, 민감도, 특이도를 포함한 다섯 가지 평가 지표를 사용하였다.
  • 일관성 있는 일반화 능력을 확보하기 위해 환자 간 중복 없이 훈련 및 테스트 세트를 분리한 교차 데이터셋 검증을 실시하였다.
  • 모델 성능 최적화를 위해 MIMIC-III 및 eICU-CRD 데이터셋을 활용한 후, 모든 세 데이터셋에서 독립적으로 검증하였다.

실험 결과

연구 질문

  • RQ1기존 임상 점수 체계에 비해 해석 가능한 기계학습 모델이 노인 MODS 환자에서 조기 병원 사망 예측을 향상시킬 수 있는가?
  • RQ2다양한 독립적이고 다기관 중환자실 데이터베이스에서 검증했을 때 모델의 일반화 능력은 어떠한가?
  • RQ3SHAP 해석을 통해 어떤 임상적 특징이 사망 예측에 가장 크게 기여하는가?
  • RQ4다양한 출처의 데이터를 통합하면 MODS 사망 예측의 성능과 강건성이 향상되는가?

주요 결과

  • MIMIC-III 데이터셋에서 모델은 AUC 0.858, 민감도 0.834, 특이도 0.705를 기록하였다.
  • eICU-CRD 데이터셋에서 모델은 AUC 0.849, 민감도 0.763, 특이도 0.784를 기록하였다.
  • PLAGH-S 데이터셋에서 모델은 AUC 0.838, 민감도 0.882, 특이도 0.691를 기록하였다.
  • 모델은 모든 검증 데이터셋에서 AUC 측면에서 기준 모델과 일반적으로 사용되는 임상 점수 체계를 뛰어넘었다.
  • SHAP 분석을 통해 사망 위험도 특성의 해석 가능한 순위가 도출되어 임상적 신뢰도와 활용도가 향상되었다.
  • 다양하고 중복되지 않는 중환자실 데이터셋 간에 강력한 일반화 능력을 보이며, 모델의 강건성과 외부 타당성이 확인되었다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.