[论文解读] Capabilities of Gemini Models in Medicine
Med-Gemini,基于 Gemini 的面向医学的多模态模型家族,在14个医学基准中的10个达到最先进水平,在有对比时超越GPT-4,并展现出强大的长上下文和多模态能力,具备网页检索集成和自定义编码器。
Excellence in a wide variety of medical applications poses considerable challenges for AI, requiring advanced reasoning, access to up-to-date medical knowledge and understanding of complex multimodal data. Gemini models, with strong general capabilities in multimodal and long-context reasoning, offer exciting possibilities in medicine. Building on these core strengths of Gemini, we introduce Med-Gemini, a family of highly capable multimodal models that are specialized in medicine with the ability to seamlessly use web search, and that can be efficiently tailored to novel modalities using custom encoders. We evaluate Med-Gemini on 14 medical benchmarks, establishing new state-of-the-art (SoTA) performance on 10 of them, and surpass the GPT-4 model family on every benchmark where a direct comparison is viable, often by a wide margin. On the popular MedQA (USMLE) benchmark, our best-performing Med-Gemini model achieves SoTA performance of 91.1% accuracy, using a novel uncertainty-guided search strategy. On 7 multimodal benchmarks including NEJM Image Challenges and MMMU (health & medicine), Med-Gemini improves over GPT-4V by an average relative margin of 44.5%. We demonstrate the effectiveness of Med-Gemini's long-context capabilities through SoTA performance on a needle-in-a-haystack retrieval task from long de-identified health records and medical video question answering, surpassing prior bespoke methods using only in-context learning. Finally, Med-Gemini's performance suggests real-world utility by surpassing human experts on tasks such as medical text summarization, alongside demonstrations of promising potential for multimodal medical dialogue, medical research and education. Taken together, our results offer compelling evidence for Med-Gemini's potential, although further rigorous evaluation will be crucial before real-world deployment in this safety-critical domain.
研究动机与目标
- 利用 Gemini 基础提升医学 AI 的临床推理与多模态理解。
- 在文本、多模态和长上下文基准上评估 Med-Gemini,以确立最先进水平。
- 展示网页检索落地与模态特定编码器在医学任务中的好处。
- 评估在医疗笔记摘要和转诊信函生成等现实世界应用中的实用性。
- 强调安全性考量及在部署前需要严格验证。
提出的方法
- 对 Gemini 1.0 Ultra 与 1.0 Pro 进行微调,创建 Med-Gemini L 1.0(基于网页检索)与 Med-Gemini M 1.0/1.5 以用于多模态任务。
- 开发包含推理说明(CoTs)和思维链提示的自监督训练数据;创建 MedQA-R 与 MedQA-RS 数据集。
- 在推理阶段实现不确定性引导的搜索,在需要时触发网页检索并将检索结果合并到提示中。
- 在八个多模态医疗数据集上进行微调,形成 Med-Gemini-M 1.5,并为新模态(如心电图 ECG)创建专用编码器。
- 利用长上下文配置和推理链来应对“海里挑针”的电子病历检索和医疗视频理解任务。
- 在涵盖 14 个医学基准的 25 项任务上进行基准测试,包括 MedQA(USMLE)、NEJM CPC 和 GeneTuring。

实验结果
研究问题
- RQ1Med-Gemini 是否能够在广泛的医学任务(文本、多模态、长上下文)中达到最先进的结果?
- RQ2网页检索落地与不确定性引导推理是否提升医学推理的准确性?
- RQ3模态特定编码器在将 Med-Gemini 扩展到新型医学数据(如 ECG)和长上下文电子病历方面有多大帮助?
- RQ4Med-Gemini 在医学笔记摘要和转诊信函生成等实际任务中的真实世界应用价值如何?
- RQ5在存在直接对比的可比基准上,Med-Gemini 与 GPT-4 的对比结果如何?
主要发现
- Med-Gemini 在 14 项医学基准中的 10 项实现了最先进性能,在可对比的情况下超越了 GPT-4。
- 在 MedQA(USMLE)上,表现最佳的 Med-Gemini 模型达到 91.1% 的准确率,比 Med-PaLM 2 高出 4.6%。
- 不确定性引导搜索推广至 NEJM CPC 与 GeneTuring 基准,取得最先进结果。
- 在多模态基准中,Med-Gemini 在 7 项任务中有 5 项达到最先进水平,平均相对边际较 GPT-4V 提升 44.5%。
- 长上下文能力实现对“海里挑针”的电子病历理解和医疗视频问答的最先进结果。
- 除了基准测试外,Med-Gemini 在医疗笔记摘要和转诊信函生成等现实世界应用方面显示出潜在实用性,并伴随定性的多模态对话演示。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。