Skip to main content
QUICK REVIEW

[论文解读] Health AI Developer Foundations

Atilla P. Kiraly, Sebastien Baur|arXiv (Cornell University)|Nov 22, 2024
Artificial Intelligence in Healthcare and Education被引用 4
一句话总结

健康人工智能开发基础模型(HAI-DEF)推出了一套预训练的、领域特定的基础模型、工具和代码示例,用于医学人工智能开发,涵盖放射科、组织病理学、皮肤病学和音频模态。通过提供高质量、任务无关的嵌入表示,如ELIXR-C和HeAR等模型,HAI-DEF降低了数据、计算和标注需求,从而在低数据环境下实现更快、更高效的模型开发,并通过针对性微调提升了公平性。

ABSTRACT

Robust medical Machine Learning (ML) models have the potential to revolutionize healthcare by accelerating clinical research, improving workflows and outcomes, and producing novel insights or capabilities. Developing such ML models from scratch is cost prohibitive and requires substantial compute, data, and time (e.g., expert labeling). To address these challenges, we introduce Health AI Developer Foundations (HAI-DEF), a suite of pre-trained, domain-specific foundation models, tools, and recipes to accelerate building ML for health applications. The models cover various modalities and domains, including radiology (X-rays and computed tomography), histopathology, dermatological imaging, and audio. These models provide domain specific embeddings that facilitate AI development with less labeled data, shorter training times, and reduced computational costs compared to traditional approaches. In addition, we utilize a common interface and style across these models, and prioritize usability to enable developers to integrate HAI-DEF efficiently. We present model evaluations across various tasks and conclude with a discussion of their application and evaluation, covering the importance of ensuring efficacy, fairness, and equity. Finally, while HAI-DEF and specifically the foundation models lower the barrier to entry for ML in healthcare, we emphasize the importance of validation with problem- and population-specific data for each desired usage setting. This technical report will be updated over time as more modalities and features are added.

研究动机与目标

  • 解决从零开始训练稳健医学机器学习模型所面临的高昂成本、数据稀缺和计算负担问题。
  • 通过提供可访问的、预训练的基础模型和工具,降低医疗领域人工智能开发的入门门槛。
  • 通过领域特定的嵌入表示以及对问题和人群特定数据的微调支持,提升模型性能和公平性。
  • 通过提供标准化、可使用的接口和开源权重模型,加速人工智能在临床环境中的实际部署。
  • 通过支持反馈整合和向新模态与应用场景的扩展,促进社区驱动的持续改进。

提出的方法

  • HAI-DEF 采用多种在大规模多样化医学数据集上训练的基础模型,使用对比学习(SupCon)、CLIP 和 BLIP-2 技术实现图像与文本的联合表征学习。
  • 胸部X光(CXR)基础模型采用 EfficientNet-L2 主干网络,并在配对的X光图像和放射科报告上进行训练,以学习具有临床意义的嵌入表示。
  • 在音频方面,HeAR(健康声学表征)是一种在咳嗽和呼吸录音上训练的基础模型,用于提取下游任务所需的有意义嵌入。
  • 该平台提供研究用接口和开源权重模型,支持通过开源容器和专用库(如组织病理学专用库)进行部署。
  • 所有模型均采用统一接口和一致风格,以提升开发者的可用性和集成便捷性。
  • 建议在本地、任务特定的数据上进行微调,以提升性能,尤其适用于罕见疾病或代表性不足的人群。
Figure 1 : Data efficiency comparison of four of our foundation models against established approaches. CXR (upper left) shows the average performance across six binary classification tasks with the original CXR Foundation model and the new ELIXR-B model; for Pathology (upper right), it focuses on a
Figure 1 : Data efficiency comparison of four of our foundation models against established approaches. CXR (upper left) shows the average performance across six binary classification tasks with the original CXR Foundation model and the new ELIXR-B model; for Pathology (upper right), it focuses on a

实验结果

研究问题

  • RQ1预训练的基础模型是否能显著减少训练下游医学人工智能模型所需的标注数据和计算资源?
  • RQ2HAI-DEF 的领域特定嵌入在多样化的临床任务中表现如何,特别是在低数据环境下?
  • RQ3当在人群特定或设备特定数据上进行微调时,HAI-DEF 模型在多大程度上能提升公平性和公平性?
  • RQ4基于 CLIP 和 BLIP-2 的模型在零样本和少样本医学影像任务中的性能特征有何差异?
  • RQ5社区反馈和迭代模型更新在提升健康人工智能基础模型长期实用性与泛化能力方面发挥何种作用?

主要发现

  • HAI-DEF 基础模型(如 ELIXR-C 和 HeAR)在下游任务中展现出强大的零样本性能,显著减少了对大规模微调的需求。
  • 使用领域特定嵌入可实现鲁棒的模型性能,且所需标注数据显著减少,尤其在低数据环境下优势明显。
  • 在本地、人群特定数据上进行微调可提升性能并减少偏差,尤其对罕见疾病或代表性不足群体效果显著。
  • HeAR 在通过智能手机采集的新音频数据上测试时表现出强大的泛化能力,表明其具备出色的现实世界适应性。
  • 平台的标准化接口和开源权重模型可加速集成与迭代,显著缩短研究与部署周期。
  • 尽管性能优异,HAI-DEF 模型目前尚不支持图像分割或生成任务,且在设备端部署可能需要模型蒸馏。
Figure 2 : Performance of ELIXR-C, ELIXR-B, and the original CXR Foundation embeddings for data-efficient classification. The ROC AUC results of a linear probe are shown averaged across 2 datasets (CheXpert and Chest X-ray14) for seven findings: atelectasis, cardiomegaly, airspace opacity, fracture,
Figure 2 : Performance of ELIXR-C, ELIXR-B, and the original CXR Foundation embeddings for data-efficient classification. The ROC AUC results of a linear probe are shown averaged across 2 datasets (CheXpert and Chest X-ray14) for seven findings: atelectasis, cardiomegaly, airspace opacity, fracture,

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。