[Paper Review] Large Language Models in Medical Term Classification and Unexpected Misalignment Between Response and Reasoning
This study evaluates state-of-the-art large language models (LLMs) like GPT-4, GPT-3.5, Falcon, and LLaMA 2 in classifying mild cognitive impairment (MCI) from clinical discharge summaries using the MIMIC-IV v2.2 dataset. Despite GPT-4's strong reasoning and interpretability, it exhibits significant misalignment between its reasoning and final response, highlighting a critical gap in reliability for clinical deployment despite high performance.
This study assesses the ability of state-of-the-art large language models (LLMs) including GPT-3.5, GPT-4, Falcon, and LLaMA 2 to identify patients with mild cognitive impairment (MCI) from discharge summaries and examines instances where the models' responses were misaligned with their reasoning. Utilizing the MIMIC-IV v2.2 database, we focused on a cohort aged 65 and older, verifying MCI diagnoses against ICD codes and expert evaluations. The data was partitioned into training, validation, and testing sets in a 7:2:1 ratio for model fine-tuning and evaluation, with an additional metastatic cancer dataset from MIMIC III used to further assess reasoning consistency. GPT-4 demonstrated superior interpretative capabilities, particularly in response to complex prompts, yet displayed notable response-reasoning inconsistencies. In contrast, open-source models like Falcon and LLaMA 2 achieved high accuracy but lacked explanatory reasoning, underscoring the necessity for further research to optimize both performance and interpretability. The study emphasizes the significance of prompt engineering and the need for further exploration into the unexpected reasoning-response misalignment observed in GPT-4. The results underscore the promise of incorporating LLMs into healthcare diagnostics, contingent upon methodological advancements to ensure accuracy and clinical coherence of AI-generated outputs, thereby improving the trustworthiness of LLMs for medical decision-making.
Motivation & Objective
- To assess the performance of state-of-the-art LLMs in identifying mild cognitive impairment (MCI) from clinical discharge summaries.
- To investigate the consistency between LLM-generated reasoning and final classification responses.
- To compare the interpretability and accuracy of proprietary models (e.g., GPT-4) versus open-source models (e.g., Falcon, LLaMA 2).
- To evaluate the impact of prompt engineering on model behavior and clinical coherence.
- To identify methodological challenges in deploying LLMs for reliable medical diagnostics.
Proposed method
- Fine-tuned LLMs on a 7:2:1 train/validation/test split of MIMIC-IV v2.2 data from patients aged 65+ with confirmed MCI via ICD codes and expert review.
- Used complex, multi-step prompts to elicit detailed reasoning from models before classification.
- Evaluated model responses against gold-standard MCI diagnoses to detect reasoning-response misalignment.
- Conducted ablation studies using a metastatic cancer dataset from MIMIC-III to test reasoning consistency across domains.
- Compared performance across proprietary (GPT-4, GPT-3.5) and open-source (Falcon, LLaMA 2) models in both accuracy and interpretability.
- Employed qualitative and quantitative analysis to measure alignment between reasoning steps and final predictions.
Experimental results
Research questions
- RQ1How accurately can LLMs classify mild cognitive impairment from clinical discharge summaries?
- RQ2To what extent do LLMs' reasoning processes align with their final classification responses?
- RQ3How do proprietary models like GPT-4 compare to open-source models like Falcon and LLaMA 2 in terms of accuracy and interpretability?
- RQ4Does prompt engineering improve the consistency and clinical coherence of LLM-generated reasoning?
- RQ5How generalizable is reasoning consistency across different medical conditions, such as MCI and metastatic cancer?
Key findings
- GPT-4 demonstrated superior interpretative capabilities, particularly in complex prompting scenarios, but exhibited notable misalignment between reasoning and final response.
- Open-source models like Falcon and LLaMA 2 achieved high classification accuracy but lacked detailed or consistent reasoning, limiting clinical trust.
- Despite strong performance, GPT-4’s reasoning often did not logically support its final prediction, indicating a critical reliability gap.
- The study identified a recurring pattern of reasoning-response misalignment in GPT-4, especially under complex or ambiguous clinical prompts.
- Performance on the metastatic cancer subset confirmed that reasoning inconsistencies persisted across different medical conditions, suggesting a systemic issue.
- The results emphasize that high accuracy alone is insufficient for clinical deployment; reasoning coherence and interpretability are essential for trustworthy AI in healthcare.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.