[Paper Review] Evaluation of GPT-3.5 and GPT-4 for supporting real-world information needs in healthcare delivery
The study assesses GPT-3.5 and GPT-4 as informatics consultation aids for physician questions in healthcare, finding limited concordance with expert reports and no majority indications of harm, highlighting the need for prompt engineering and model tailoring.
Despite growing interest in using large language models (LLMs) in healthcare, current explorations do not assess the real-world utility and safety of LLMs in clinical settings. Our objective was to determine whether two LLMs can serve information needs submitted by physicians as questions to an informatics consultation service in a safe and concordant manner. Sixty six questions from an informatics consult service were submitted to GPT-3.5 and GPT-4 via simple prompts. 12 physicians assessed the LLM responses' possibility of patient harm and concordance with existing reports from an informatics consultation service. Physician assessments were summarized based on majority vote. For no questions did a majority of physicians deem either LLM response as harmful. For GPT-3.5, responses to 8 questions were concordant with the informatics consult report, 20 discordant, and 9 were unable to be assessed. There were 29 responses with no majority on "Agree", "Disagree", and "Unable to assess". For GPT-4, responses to 13 questions were concordant, 15 discordant, and 3 were unable to be assessed. There were 35 responses with no majority. Responses from both LLMs were largely devoid of overt harm, but less than 20% of the responses agreed with an answer from an informatics consultation service, responses contained hallucinated references, and physicians were divided on what constitutes harm. These results suggest that while general purpose LLMs are able to provide safe and credible responses, they often do not meet the specific information need of a given question. A definitive evaluation of the usefulness of LLMs in healthcare settings will likely require additional research on prompt engineering, calibration, and custom-tailoring of general purpose models.
Motivation & Objective
- Assess whether two large language models (GPT-3.5 and GPT-4) can safely answer physician-submitted informatics questions.
- Evaluate concordance between LLM responses and established informatics consultation reports.
- Identify safety concerns, including potential patient harm and hallucinations, in real-world clinical queries.
Proposed method
- Submit 66 physician questions from an informatics consultation service to GPT-3.5 and GPT-4 with simple prompts.
- Have 12 physicians assess whether LLM responses pose patient harm and whether they concord with the informatics consultation reports.
- Summarize physician assessments by majority vote to determine safety and concordance.
- Report counts for concordant, discordant, and unable to assess for each model.
Experimental results
Research questions
- RQ1Can GPT-3.5 and GPT-4 provide safe responses to real-world physician information needs in healthcare delivery?
- RQ2To what extent are LLM responses concordant with established informatics consultation reports?
- RQ3What are the observed patterns of harm, hallucination, or misalignment in LLM outputs for clinical queries?
Key findings
- No majority of physicians deemed any LLM response harmful for any question.
- GPT-3.5: 8 concordant, 20 discordant, 9 unable to assess; 29 with no majority across Agree/Disagree/Unable.
- GPT-4: 13 concordant, 15 discordant, 3 unable to assess; 35 with no majority across Agree/Disagree/Unable.
- Responses from both LLMs largely lacked overt harm but contained hallucinated references and often did not align with the informatics consultation reports.
- Less than 20% of responses agreed with the informatics consultation service answer.
- Indicates general-purpose LLMs can be safe but not reliably useful without further prompt engineering and customization.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.