[Paper Review] Medical Misinformation in AI-Assisted Self-Diagnosis: Development of a Method (EvalPrompt) for Analyzing Large Language Models
This paper introduces EvalPrompt, a novel method to evaluate large language models (LLMs) like ChatGPT in realistic self-diagnosis scenarios by testing open-ended medical questions without answer options. It reveals that LLMs perform significantly worse than previously assumed, exhibit overconfidence in incorrect recommendations, and fail to include disclaimers or indicate uncertainty—amplifying medical misinformation risks in real-world use.
Rapid integration of large language models (LLMs) in health care is sparking global discussion about their potential to revolutionize health care quality and accessibility. At a time when improving health care quality and access remains a critical concern for countries worldwide, the ability of these models to pass medical examinations is often cited as a reason to use them for medical training and diagnosis. However, the impact of their inevitable use as a self-diagnostic tool and their role in spreading healthcare misinformation has not been evaluated. This study aims to assess the effectiveness of LLMs, particularly ChatGPT, from the perspective of an individual self-diagnosing to better understand the clarity, correctness, and robustness of the models. We propose the comprehensive testing methodology evaluation of LLM prompts (EvalPrompt). This evaluation methodology uses multiple-choice medical licensing examination questions to evaluate LLM responses. We use open-ended questions to mimic real-world self-diagnosis use cases, and perform sentence dropout to mimic realistic self-diagnosis with missing information. Human evaluators then assess the responses returned by ChatGPT for both experiments for clarity, correctness, and robustness. The results highlight the modest capabilities of LLMs, as their responses are often unclear and inaccurate. As a result, medical advice by LLMs should be cautiously approached. However, evidence suggests that LLMs are steadily improving and could potentially play a role in healthcare systems in the future. To address the issue of medical misinformation, there is a pressing need for the development of a comprehensive self-diagnosis dataset. This dataset could enhance the reliability of LLMs in medical applications by featuring more realistic prompt styles with minimal information across a broader range of medical fields.
Motivation & Objective
- To critically assess the real-world performance of LLMs like ChatGPT in self-diagnosis, moving beyond multiple-choice exam benchmarks.
- To investigate how LLMs contribute to medical misinformation when used by non-experts for open-ended symptom queries.
- To develop a repeatable, human-assessor-based evaluation framework that captures nuanced response quality beyond simple correctness.
- To examine whether LLMs can recognize uncertainty or validate their own responses, especially in high-risk medical contexts.
- To provide a methodological blueprint for evaluating LLMs in healthcare applications that prioritizes real-world usability over exam performance.
Proposed method
- Adapts a subset of single-answer USMLE Step 1 questions into open-ended prompts to simulate real user self-diagnosis.
- Employs non-expert human assessors to rate responses on a granular scale: Correct, Partially Correct, Incorrect, or Ambiguous.
- Conducts sensitivity analysis via sentence dropout (ablation) to test robustness of LLM responses to missing information.
- Asks ChatGPT to self-evaluate its own answers to assess its ability to detect errors or uncertainty.
- Uses aggregated human ratings to quantify performance degradation and behavioral anomalies in LLM outputs.
- Establishes a repeatable, coder-reliability-aware methodology for evaluating LLMs on open-ended medical reasoning tasks.
Experimental results
Research questions
- RQ1How does the performance of LLMs like ChatGPT on medical self-diagnosis tasks change when answer options are not provided, compared to multiple-choice settings?
- RQ2To what extent do LLMs fail to include disclaimers or indicate uncertainty in their medical responses, even when incorrect?
- RQ3How does information dropout (removal of key clinical details) affect the perceived correctness of LLM-generated medical advice?
- RQ4Can LLMs accurately self-evaluate their own responses, particularly in identifying incorrect or overconfident answers?
- RQ5What behavioral patterns—such as overconfidence or unnecessary caution—emerge in LLM responses during medical reasoning tasks?
Key findings
- LLM performance drops significantly when answer options are not provided, indicating that multiple-choice benchmarks overstate real-world diagnostic capability.
- ChatGPT frequently fails to include disclaimers or indicate uncertainty, even when recommending incorrect or potentially harmful medical actions.
- The model exhibits overconfidence in incorrect recommendations, increasing the risk of spreading medical misinformation in self-diagnosis contexts.
- ChatGPT often categorizes medically correct answers as 'Partially Correct' when they include additional, non-essential but clinically appropriate advice, indicating a lack of clinical intuition.
- When asked to self-evaluate, ChatGPT shows poor awareness of its own errors, suggesting it cannot reliably serve as a fact-checking tool.
- Even minor information dropout (e.g., removing a single clinical sentence) significantly alters the perceived correctness of responses, highlighting fragility in reasoning under real-world data variation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.