Skip to main content
QUICK REVIEW

[Paper Review] AI-Driven CT-based quantification, staging and short-term outcome prediction of COVID-19 pneumonia

Guillaume Chassagnon, Maria Vakalopoulou|arXiv (Cornell University)|Apr 20, 2020
COVID-19 diagnosis using AI24 references17 citations
TL;DR

This study presents an AI-driven framework for automated quantification, staging, and short-term outcome prediction of COVID-19 pneumonia using chest CT scans. It integrates radiomics features with multiple machine learning classifiers, achieving high accuracy (up to 0.81) in predicting severe outcomes and intubation/death, demonstrating the potential of AI to support clinical decision-making in critical care settings.

ABSTRACT

Chest computed tomography (CT) is widely used for the management of Coronavirus disease 2019 (COVID-19) pneumonia because of its availability and rapidity. The standard of reference for confirming COVID-19 relies on microbiological tests but these tests might not be available in an emergency setting and their results are not immediately available, contrary to CT. In addition to its role for early diagnosis, CT has a prognostic role by allowing visually evaluating the extent of COVID-19 lung abnormalities. The objective of this study is to address prediction of short-term outcomes, especially need for mechanical ventilation. In this multi-centric study, we propose an end-to-end artificial intelligence solution for automatic quantification and prognosis assessment by combining automatic CT delineation of lung disease meeting performance of experts and data-driven identification of biomarkers for its prognosis. AI-driven combination of variables with CT-based biomarkers offers perspectives for optimal patient management given the shortage of intensive care beds and ventilators.

Motivation & Objective

  • To develop an automated AI system for quantifying and staging COVID-19 pneumonia from chest CT scans.
  • To predict short-term clinical outcomes, including disease severity and need for intubation or death.
  • To evaluate the performance of multiple machine learning classifiers in distinguishing between patient outcome groups.
  • To identify the most predictive radiomics features from CT images for clinical outcome prediction.
  • To validate the robustness of the model using independent test sets and cross-validation.

Proposed method

  • The framework extracts 12 radiomics features from lung regions in contrast-enhanced CT scans, including intensity, texture, and shape-based metrics.
  • Features are derived from Gray-Level Size Zone Matrix (GLSZM) and Gray-Level Run-Length Matrix (GLRLM) to capture texture heterogeneity.
  • Multiple classifiers—SVMs (linear, polynomial, RBF), decision trees, random forests, AdaBoost, Naive Bayes, and ensemble models—are trained and evaluated.
  • Model performance is assessed using balanced accuracy, weighted precision, sensitivity, and specificity on training and independent test sets.
  • A hierarchical classification strategy is employed to predict both disease severity and critical outcomes (intubation or death).
  • Statistical correlation analysis is performed between clinical outcomes and radiomics features to identify significant predictors.

Experimental results

Research questions

  • RQ1Can radiomics features extracted from CT scans reliably predict the severity of COVID-19 pneumonia?
  • RQ2Which machine learning models achieve the highest performance in predicting short-term clinical outcomes from CT data?
  • RQ3Which radiomics features show the strongest correlation with severe outcomes or mortality?
  • RQ4How does the performance of ensemble models compare to individual classifiers in outcome prediction?
  • RQ5Can the model generalize well to independent test sets, indicating clinical robustness?

Key findings

  • The ensemble classifier achieved a balanced accuracy of 0.81 (±0.01) in predicting severe vs. non-severe outcomes on the test set.
  • For intubation or death prediction, the ensemble model reached a balanced accuracy of 0.81 (±0.01), with 0.88 sensitivity and 0.74 specificity.
  • The RBF-SVM classifier showed the highest training accuracy (0.90) for intubation/death prediction but poor generalization (0.62 test accuracy).
  • The most predictive features included volume, maximum attenuation, and non-uniformity on GLSZM and GLRLM matrices, with correlation coefficients up to -0.3971 (volume vs. severe outcome).
  • The Gaussian Naive Bayes and L-SVM classifiers demonstrated stable performance across both outcome tasks, with balanced accuracy above 0.70 on test sets.
  • The hierarchical classifier design improved model interpretability and allowed for sequential prediction of severity and critical outcomes.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.