[论文解读] Evaluation of ChatGPT-Generated Medical Responses: A Systematic Review and Meta-Analysis
本篇系统性综述与荟萃分析评估了ChatGPT在医学问答任务中的表现,共分析了从3520篇文献中筛选出的17项研究。结果显示,其总体综合准确率为56%(95%置信区间:51%–60%,I² = 87%),凸显了研究间显著的异质性及报告不一致的问题,限制了结果的可靠性,呼吁未来研究建立标准化、透明的评估框架。
Large language models such as ChatGPT are increasingly explored in medical domains. However, the absence of standard guidelines for performance evaluation has led to methodological inconsistencies. This study aims to summarize the available evidence on evaluating ChatGPT's performance in medicine and provide direction for future research. We searched ten medical literature databases on June 15, 2023, using the keyword "ChatGPT". A total of 3520 articles were identified, of which 60 were reviewed and summarized in this paper and 17 were included in the meta-analysis. The analysis showed that ChatGPT displayed an overall integrated accuracy of 56% (95% CI: 51%-60%, I2 = 87%) in addressing medical queries. However, the studies varied in question resource, question-asking process, and evaluation metrics. Moreover, many studies failed to report methodological details, including the version of ChatGPT and whether each question was used independently or repeatedly. Our findings revealed that although ChatGPT demonstrated considerable potential for application in healthcare, the heterogeneity of the studies and insufficient reporting may affect the reliability of these results. Further well-designed studies with comprehensive and transparent reporting are needed to evaluate ChatGPT's performance in medicine.
研究动机与目标
- 总结现有关于ChatGPT在医学情境中表现的证据。
- 识别当前评估ChatGPT用于医学问答研究中的方法学不一致性。
- 评估各研究中报告的表现指标的可靠性和有效性。
- 基于报告和研究设计中的缺口,为未来研究提出建议。
- 为大型语言模型在医疗保健领域的标准化评估奠定基础。
提出的方法
- 截至2023年6月15日,使用关键词'ChatGPT'在十个医学数据库中进行系统性文献检索。
- 根据预设的纳入标准对研究进行筛选与选择,共审查60项研究,其中17项纳入荟萃分析。
- 采用随机效应模型进行荟萃分析,以估计各研究间的合并准确率,置信区间为95%,并使用I²统计量评估异质性。
- 评估方法学质量,包括ChatGPT版本的报告情况、问题是否重复使用以及问题来源。
- 综合分析各研究在准确率、一致性和透明度方面的发现。
- 目标发表期刊为《Journal of Biomedical Informatics》,最终结果发表于2024年卷151。
实验结果
研究问题
- RQ1在现有研究中,ChatGPT在回答医学问题方面的总体准确率是多少?
- RQ2在评估医学领域ChatGPT的研究中,其评估方法的一致性如何?
- RQ3方法学报告中的缺口在多大程度上影响了表现估计的可靠性?
- RQ4各研究在问题来源、提示策略和评估指标方面存在哪些关键差异?
- RQ5可提出哪些建议以改善未来评估工作的标准化与透明度?
主要发现
- 在纳入的研究中,ChatGPT在回答医学问题方面的总体综合准确率为56%(95%置信区间:51%–60%)。
- 研究间观察到显著的异质性,I²统计量为87%,表明结果存在显著不一致性。
- 许多研究未能报告关键的方法学细节,包括所使用的ChatGPT具体版本,以及问题是否被重复使用或独立处理。
- 报告质量参差不齐,评估方案和数据收集流程缺乏透明度。
- 尽管潜力可观,但当前证据基础因方法学变异性及报告质量差而受损,降低了对聚合表现估计结果的信心。
- 研究结果强调了在临床环境中对大型语言模型建立标准化、透明且设计良好的评估框架的迫切需求。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。