Skip to main content
QUICK REVIEW

[논문 리뷰] Foundational Models in Medical Imaging: A Comprehensive Survey and Future Vision

Bobby Azad, Reza Azad|arXiv (Cornell University)|2023. 10. 28.
Radiomics and Machine Learning in Medical Imaging인용 수 40
한 줄 요약

본 설문은 의학 영상에서의 foundation models를 검토하고, 분류 체계를 제안하며, 학습, prompting, 및 멀티모달 적응에 대해 논의하고, 도전과제와 향후 방향을 요약합니다.

ABSTRACT

Foundation models, large-scale, pre-trained deep-learning models adapted to a wide range of downstream tasks have gained significant interest lately in various deep-learning problems undergoing a paradigm shift with the rise of these models. Trained on large-scale dataset to bridge the gap between different modalities, foundation models facilitate contextual reasoning, generalization, and prompt capabilities at test time. The predictions of these models can be adjusted for new tasks by augmenting the model input with task-specific hints called prompts without requiring extensive labeled data and retraining. Capitalizing on the advances in computer vision, medical imaging has also marked a growing interest in these models. To assist researchers in navigating this direction, this survey intends to provide a comprehensive overview of foundation models in the domain of medical imaging. Specifically, we initiate our exploration by providing an exposition of the fundamental concepts forming the basis of foundation models. Subsequently, we offer a methodical taxonomy of foundation models within the medical domain, proposing a classification system primarily structured around training strategies, while also incorporating additional facets such as application domains, imaging modalities, specific organs of interest, and the algorithms integral to these models. Furthermore, we emphasize the practical use case of some selected approaches and then discuss the opportunities, applications, and future directions of these large-scale pre-trained models, for analyzing medical images. In the same vein, we address the prevailing challenges and research pathways associated with foundational models in medical imaging. These encompass the areas of interpretability, data management, computational requirements, and the nuanced issue of contextual comprehension.

연구 동기 및 목표

  • 의학 영상에서의 foundational models(FMs)과 그 핵심 개념에 대한 구조화된 개요를 제공합니다.
  • 학습 전략과 모달리티를 기반으로 의학 FMs의 분류 체계를 소개합니다.
  • 영상 모달리티와 장기에 걸친 응용을 분석하여 강점과 한계를 부각합니다.
  • 해석 가능성, 데이터 프라이버시, 계산 비용, 도메인 지식 통합과 같은 도전과제를 논의합니다.
  • 임상 실무에서의 MFMs를 위한 향후 방향과 연구 경로를 제안합니다.

제안 방법

  • foundation models를 정의하고 사전 학습 목표(대조적, 생성적, 하이브리드)를 요약합니다.
  • 의학 FMs를 여섯 그룹으로 분류합니다: VPM-Generalist, TPM-Hybrid, TPM-Contrastive, TPM-Generative, VPM-Adaptations, TPM-Conversational.
  • 의료 영상 작업에 초점을 맞추어 텍스트 기반 프롬프트 모델과 시각 프롬프트 모델의 대표적 연구를 검토합니다.
  • 프롬프트 엔지니어링, 지시 정렬(instruction-aligning), 그리고 LM에서 영감을 받은 프롬 prompting이 비전-언어 의료 작업으로 어떻게 확장되는지 논의합니다.
  • 오픈 소스 구현을 집계하고 지속 업데이트를 위한 참조를 제공합니다(GitHub 이니셔티브).
(a) Algorithms
(a) Algorithms

실험 결과

연구 질문

  • RQ1의학 기반 모델을 가능하게 하는 근본 개념과 학습 전략은 무엇인가?
  • RQ2의학 FMs를 프롬프트 및 모달리티에 따라 어떻게 분류할 수 있으며, 각 범주별 대표적 접근법은 무엇인가?
  • RQ3MFMs의 주요 임상 기회, 한계 및 도전 과제(해석 가능성, 프라이버시, 계산, 분포 이동)는 무엇인가?
  • RQ4더 넓은 임상 영향력을 위해 MFMs를 이끌 미래 방향과 미해결 연구 질문은 무엇인가?

주요 결과

  • FMs는 다중 모달 해석과 한정된 라벨 데이터로의 맥락 내 적응을 가능하게 한다.
  • 의학 FMs를 다루는 여섯 그룹 분류 체계는 텍스트 기반 및 시각 기반 프롬프트를 포함한 대화형 및 일반ist 변형을 포괄한다.
  • 주요 예시들(MedCLIP, BiomedCLIP, MI-Zero, BioViL-T, Med-Flamingo)은 방사선학, 병리학 및 기타 도메인에서 제로샷, 파샷 및 다중 모달 기능을 보여준다.
  • 프롬프트링, 지시 정렬, 도메인 지식 통합은 의료 영상 FM의 성능 향상에 핵심이다.
  • 의료 환경에서 데이터 프라이버시 보존과 연합 학습은 주요 이점으로 강조된다.
  • 이 설문은 지속적인 업데이트와 빠른 확산 및 평가를 위한 오픈 소스 자원을 강조한다.
(b) Modalities
(b) Modalities

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.