Skip to main content
QUICK REVIEW

[Paper Review] Multimodal Neurodegenerative Disease Subtyping Explained by ChatGPT

Diego Machado Reyes, Hanqing Chao|arXiv (Cornell University)|Jan 31, 2024
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This paper proposes a multimodal deep learning framework, Tri-COAT, that integrates baseline imaging, genetic, and clinical data to subtype Alzheimer’s disease at early (prodromal) stages, achieving improved classification over single-modality models. It leverages large language models like ChatGPT to generate clinically interpretable explanations of cross-modal feature associations, revealing biologically plausible links between brain atrophy, genetic risk, and cognitive decline.

ABSTRACT

Alzheimer's disease (AD) is the most prevalent neurodegenerative disease; yet its currently available treatments are limited to stopping disease progression. Moreover, effectiveness of these treatments is not guaranteed due to the heterogenetiy of the disease. Therefore, it is essential to be able to identify the disease subtypes at a very early stage. Current data driven approaches are able to classify the subtypes at later stages of AD or related disorders, but struggle when predicting at the asymptomatic or prodromal stage. Moreover, most existing models either lack explainability behind the classification or only use a single modality for the assessment, limiting scope of its analysis. Thus, we propose a multimodal framework that uses early-stage indicators such as imaging, genetics and clinical assessments to classify AD patients into subtypes at early stages. Similarly, we build prompts and use large language models, such as ChatGPT, to interpret the findings of our model. In our framework, we propose a tri-modal co-attention mechanism (Tri-COAT) to explicitly learn the cross-modal feature associations. Our proposed model outperforms baseline models and provides insight into key cross-modal feature associations supported by known biological mechanisms.

Motivation & Objective

  • Address the challenge of early Alzheimer’s disease subtyping due to disease heterogeneity and limited early biomarkers.
  • Overcome the limitations of single-modality models and non-explainable deep learning approaches in classifying prodromal AD.
  • Develop a multimodal fusion framework that leverages early-stage indicators (imaging, genetics, clinical) for improved subtyping accuracy.
  • Enhance model interpretability by integrating large language models (e.g., ChatGPT) to generate clinically meaningful explanations of feature-level associations.
  • Establish a foundation for personalized medicine in neurodegenerative diseases through explainable, multimodal subtyping.

Proposed method

  • Propose a tri-modal co-attention mechanism (Tri-COAT) to explicitly model cross-modal feature interactions across imaging, genetic, and clinical data.
  • Use early-stage baseline data—without longitudinal follow-up—to train the model for subtyping at the prodromal or asymptomatic stage.
  • Apply gradient-based attribution methods (e.g., Integrated Gradients) to identify salient features driving model predictions.
  • Design structured prompts for large language models (e.g., ChatGPT, Bard) to interpret top-10 attributions, including feature names, patient values, and deviations from population norms.
  • Generate natural language reports that link each salient feature to known biological mechanisms (e.g., brain atrophy in the right insula linked to rapid AD progression).
  • Validate LLM-generated explanations against clinical literature to assess accuracy and biological plausibility, while acknowledging potential hallucinations.
Figure 1: The three main multimodal fusion strategies, early, intermediate and late fusion, for deep learning methods.
Figure 1: The three main multimodal fusion strategies, early, intermediate and late fusion, for deep learning methods.

Experimental results

Research questions

  • RQ1Can a multimodal deep learning model effectively subtype Alzheimer’s disease using only baseline (non-longitudinal) data from imaging, genetics, and clinical assessments?
  • RQ2How do cross-modal feature associations identified by the model align with known biological mechanisms of Alzheimer’s disease?
  • RQ3To what extent can large language models like ChatGPT generate accurate, clinically relevant explanations of model predictions and feature attributions?
  • RQ4Can LLM-generated interpretations improve the transparency and trustworthiness of multimodal deep learning models in neurodegenerative disease subtyping?
  • RQ5Does the model’s ability to detect atypical or protective features (e.g., higher cortical thickness in the cuneus) reflect known compensatory mechanisms in AD?

Key findings

  • The Tri-COAT model outperformed baseline models in classifying Alzheimer’s disease subtypes using only baseline multimodal data, demonstrating improved classification performance.
  • The model identified key cross-modal biomarker networks, such as reduced volume in the right insula being linked to rapid disease progression, which aligns with known neuroanatomical and pathological findings.
  • Large language models (ChatGPT and Bard) generated explanations that were consistent with clinical literature, such as interpreting higher cortical thickness in the right cuneus as a potential compensatory mechanism.
  • LLM-generated reports provided biologically plausible interpretations of salient features, including the functional roles of brain regions and the implications of abnormal measurements relative to population norms.
  • The model detected and explained atypical patterns—such as a higher-than-average cortical thickness in a region typically atrophied in AD—demonstrating sensitivity to protective or compensatory features.
  • Despite promising results, LLMs occasionally produced hallucinated or speculative claims, highlighting a key limitation for clinical deployment.
Figure 2: Illustration of the proposed framework for AD subtyping, consisting of two main sections: (a) single modality encoding and (b) tri-modal attention and joint encoding.
Figure 2: Illustration of the proposed framework for AD subtyping, consisting of two main sections: (a) single modality encoding and (b) tri-modal attention and joint encoding.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.