[论文解读] Clinical Validation of Medical-based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-based Evaluation
本文在眼科患者问答上对四种小型医学大模型进行实证评估,并将基于大模型的评估与临床医生分级进行对比,以评估安全性、共识与临床深度。
Domain specific large language models are increasingly used to support patient education, triage, and clinical decision making in ophthalmology, making rigorous evaluation essential to ensure safety and accuracy. This study evaluated four small medical LLMs Meerkat-7B, BioMistral-7B, OpenBioLLM-8B, and MedLLaMA3-v20 in answering ophthalmology related patient queries and assessed the feasibility of LLM based evaluation against clinician grading. In this cross sectional study, 180 ophthalmology patient queries were answered by each model, generating 2160 responses. Models were selected for parameter sizes under 10 billion to enable resource efficient deployment. Responses were evaluated by three ophthalmologists of differing seniority and by GPT-4-Turbo using the S.C.O.R.E. framework assessing safety, consensus and context, objectivity, reproducibility, and explainability, with ratings assigned on a five point Likert scale. Agreement between LLM and clinician grading was assessed using Spearman rank correlation, Kendall tau statistics, and kernel density estimate analyses. Meerkat-7B achieved the highest performance with mean scores of 3.44 from Senior Consultants, 4.08 from Consultants, and 4.18 from Residents. MedLLaMA3-v20 performed poorest, with 25.5 percent of responses containing hallucinations or clinically misleading content, including fabricated terminology. GPT-4-Turbo grading showed strong alignment with clinician assessments overall, with Spearman rho of 0.80 and Kendall tau of 0.67, though Senior Consultants graded more conservatively. Overall, medical LLMs demonstrated potential for safe ophthalmic question answering, but gaps remained in clinical depth and consensus, supporting the feasibility of LLM based evaluation for large scale benchmarking and the need for hybrid automated and clinician review frameworks to guide safe clinical deployment.
研究动机与目标
- 评估眼科领域特定大模型聊天机器人在安全性、准确性及临床实用性上的表现。
- 评估以大模型为基础的评估作为对临床评估的基准工具的可行性。
- 识别模型回答在临床深度、共识与潜在幻觉方面的差距。
- 通过选择参数在10B以下的模型,探索资源高效的部署方式。
提出的方法
- 对每个模型回答180个眼科患者提问,生成2160个回答。
- 请三位眼科医生(不同资历层级)与GPT-4-Turbo对回答进行评分,采用S.C.O.R.E.框架,基于5分Likert量表。
- 使用Spearman等级相关、Kendall tau以及核密度估计来比较LLM与临床医生分级的吻合度。
- 选择参数规模在100亿以下的模型,以实现资源高效部署。
实验结果
研究问题
- RQ1小型医学大模型是否能提供安全且临床有用的眼科解答?
- RQ2基于大模型的评估在眼科问题上与临床医生分级的对齐程度如何?
- RQ3各模型在幻觉或临床上具有误导性内容的发生率如何?
- RQ4基于大模型的基准评估是否可用于眼科医学聊天机器人的大规模评估?
主要发现
- Meerkat-7B 在资深顾问、顾问和住院医师三个层面上均获得最高的平均分(分别为3.44、4.08、4.18)。
- MedLLaMA3-v20 的表现最差,有25.5%的回答包含幻觉或临床上误导内容。
- GPT-4-Turbo的分级与临床医生评估高度一致(Spearman 0.80,Kendall 0.67),但资深顾问的分级偏保守。
- 医学大模型在眼科问答中显示出安全性潜力,但在临床深度和共识方面仍存在差距。
- 基于大模型的评估在大规模基准测试中具有可行性,支持自动化+临床评审的混合框架。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。