[Paper Review] Development and Validation of a Deep Learning Model for Prediction of Severe Outcomes in Suspected COVID-19 Infection
This study develops and validates a deep learning model, CO-RISK, that integrates electronic health record (EHR) data and chest X-ray (CXR) images to predict severe outcomes in suspected COVID-19 patients within 24–72 hours of emergency department (ED) presentation. The model achieves an AUC of 0.95 at 24 hours and 0.92 at 73 hours, outperforming clinical risk scores and physician judgment in ICU/floor triage decisions.
COVID-19 patient triaging with predictive outcome of the patients upon first present to emergency department (ED) is crucial for improving patient prognosis, as well as better hospital resources management and cross-infection control. We trained a deep feature fusion model to predict patient outcomes, where the model inputs were EHR data including demographic information, co-morbidities, vital signs and laboratory measurements, plus patient's CXR images. The model output was patient outcomes defined as the most insensitive oxygen therapy required. For patients without CXR images, we employed Random Forest method for the prediction. Predictive risk scores for COVID-19 severe outcomes ("CO-RISK" score) were derived from model output and evaluated on the testing dataset, as well as compared to human performance. The study's dataset (the "MGB COVID Cohort") was constructed from all patients presenting to the Mass General Brigham (MGB) healthcare system from March 1st to June 1st, 2020. ED visits with incomplete or erroneous data were excluded. Patients with no test order for COVID or confirmed negative test results were excluded. Patients under the age of 15 were also excluded. Finally, electronic health record (EHR) data from a total of 11060 COVID-19 confirmed or suspected patients were used in this study. Chest X-ray (CXR) images were also collected from each patient if available. Results show that CO-RISK score achieved area under the Curve (AUC) of predicting MV/death (i.e. severe outcomes) in 24 hours of 0.95, and 0.92 in 72 hours on the testing dataset. The model shows superior performance to the commonly used risk scores in ED (CURB-65 and MEWS). Comparing with physician's decisions, CO-RISK score has demonstrated superior performance to human in making ICU/floor decisions.
Motivation & Objective
- To improve early triage of suspected COVID-19 patients in the emergency department by predicting severe outcomes such as mechanical ventilation or death.
- To integrate multimodal data—EHR (demographics, comorbidities, vital signs, lab results) and CXR images—into a unified predictive model.
- To develop a clinically actionable risk score (CO-RISK) that supports resource allocation and infection control in healthcare settings.
- To validate the model’s performance against established clinical risk scores (CURB-65, MEWS) and physician decisions.
- To assess model robustness in patients without available CXR images using a Random Forest alternative.
Proposed method
- A deep feature fusion model was trained on EHR and CXR data to predict the most intensive oxygen therapy required as a proxy for severe outcomes.
- For patients without CXR images, a Random Forest model was used as a fallback to maintain prediction capability.
- The model was trained on a cohort of 11,060 confirmed or suspected COVID-19 patients from the Mass General Brigham healthcare system (March 1–June 1, 2020).
- CXR images were processed using convolutional neural network (CNN) architectures to extract radiological features, which were fused with tabular EHR features.
- The final CO-RISK score was derived from model outputs and calibrated for clinical interpretability and risk stratification.
- Model performance was evaluated using area under the receiver operating characteristic curve (AUC) on a held-out test set.
Experimental results
Research questions
- RQ1Can a deep learning model that fuses EHR and CXR data predict severe COVID-19 outcomes (e.g., mechanical ventilation or death) with higher accuracy than existing clinical risk scores?
- RQ2How does the performance of the CO-RISK model compare to that of emergency department physicians in making ICU or inpatient admission decisions?
- RQ3What is the predictive performance of the model when CXR images are unavailable, and how does the fallback Random Forest model perform?
- RQ4Does the inclusion of imaging data significantly improve prediction accuracy compared to EHR-only models?
- RQ5Can the CO-RISK score be reliably used for early triage to optimize hospital resource allocation and infection control?
Key findings
- The CO-RISK model achieved an AUC of 0.95 for predicting severe outcomes (MV/death) within 24 hours of ED presentation on the test dataset.
- At 72 hours, the model maintained strong performance with an AUC of 0.92, indicating sustained predictive accuracy over time.
- The model significantly outperformed the CURB-65 and MEWS clinical risk scores in predicting severe outcomes.
- CO-RISK demonstrated superior performance compared to physician triage decisions in distinguishing patients requiring ICU admission versus inpatient care.
- The model maintained high performance even in patients without available CXR images, using the Random Forest fallback method.
- The study confirms that multimodal integration of EHR and imaging data enhances early prediction of severe COVID-19 outcomes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.