[논문 리뷰] MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
MedTrinity-25M를 소개한다. 25M+의 다중 모달 의료 데이터셋으로 10개 모달리티와 65+ 질병에 걸친 이미지-ROI-설명 트리플렛과 다중 계층 주석이 포함되며, 자동화 파이프라인을 통해 paired text 없이 생성되었고 전문가 grounding, RAG, 및 MLLMs를 활용했다.
This paper introduces MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities with multigranular annotations for more than 65 diseases. These multigranular annotations encompass both global information, such as modality and organ detection, and local information like ROI analysis, lesion texture, and region-wise correlations. Unlike the existing multimodal datasets, which are limited by the availability of image-text pairs, we have developed the first automated pipeline that scales up multimodal data by generating multigranular visual and textual annotations in the form of image-ROI-description triplets without the need for any paired text descriptions. Specifically, data from over 30 different sources have been collected, preprocessed, and grounded using domain-specific expert models to identify ROIs related to abnormal regions. We then build a comprehensive knowledge base and prompt multimodal large language models to perform retrieval-augmented generation with the identified ROIs as guidance, resulting in multigranular textual descriptions. Compared to existing datasets, MedTrinity-25M provides the most enriched annotations, supporting a comprehensive range of multimodal tasks such as captioning and report generation, as well as vision-centric tasks like classification and segmentation. We propose LLaVA-Tri by pretraining LLaVA on MedTrinity-25M, achieving state-of-the-art performance on VQA-RAD, SLAKE, and PathVQA, surpassing representative SOTA multimodal large language models. Furthermore, MedTrinity-25M can also be utilized to support large-scale pre-training of multimodal medical AI models, contributing to the development of future foundation models in the medical domain. We will make our dataset available.
연구 동기 및 목표
- 로컬 ROI를 글로벌 질병 맥락과 연결하는 다중 계층 의료 시각 설명의 필요성을 제시한다.
- 비Paired 의료 이미지로부터 풍부한 이미지-ROI-설명 주석을 생성하는 확장 가능한 자동 파이프라인을 제공한다.
- 의료 AI 모델의 광범위한 다중모달 작업(자막 생성, 보고서 생성, 분류, 세분화) 및 대규모 사전학습을 가능하게 한다.
제안 방법
- 10개 모달리티와 65+ 질병에 걸친 25M+ 샘플에서 이미지-ROI-설명 트리플렛을 90+ 온라인 자원에서 수집한다.
- 전문가 grounding 모델을 사용해 ROI를 위치시키고 필요시 마스크를 바운딩 박스로 변환한다.
- PubMed, StatPearls, 교과서를 바탕으로 의료 지식 베이스를 구축하고 Faiss로 인덱싱해 Retrieval-Augmented Generation을 수행한다.
- 프롬프트를 통해 의료 LLM 스택(GPT-4V 부분집합 → LLaVA-Med Captioner, LLAMA3 및 다중 스케일 기능으로 강화)을 활용해 대략적인 자막, ROI, 검색된 지식에 의해 안내된 다중 계층 텍스트 설명을 생성한다.
- MedTrinity-25M에서 LLaVA-Med++를 미세 조정해 전체 25M 이미지-ROI-설명 트리플렛을 생성한다.
실험 결과
연구 질문
- RQ1비Paired 의료 이미지를 자동 grounding, Retrieval-Augmented Generation 및 MLLMs를 사용해 고품질의 다중 계층 이미지-ROI-설명 트리플렛으로 변환할 수 있는가?
- RQ2다중 계층 주석이 VQA 및 보고서 생성과 같은 하위 다중모달 의료 작업의 성능을 기존 데이터셋과 비교해 향상시키는가?
- RQ3MedTrinity-25M에서의 사전학습이 데이터셋 없이 학습된 모델에 비해 의료 VQA 벤치마크에서 우수한 결과를 낳는가?
주요 결과
- MedTrinity-25M은 90+ 소스로부터 25M개 이상의 이미지-ROI-설명 트리플렛을 포함하며 10개 모달리티와 65+ 질병에 걸친다.
- 데이터세트는 모달리티, 오르간, ROI 위치, 상호 간 관계, ROI 수준 바운딩 박스 또는 마스크를 포함한 다중 계층 텍스트 설명을 제공한다.
- GPT-4V 정렬 평가에서 SLAKE 및 MIMIC-CXR에 대해 인간 주석과의 높은 정렬을 보인다(전체 5-기준 루브릭에서 8.2/10 및 8.9/10).
- MedTrinity-25M에서 사전학습한 LLaVA-Med++는 VQA-RAD 및 PathVQA에서 최첨단 성능을 달성하고 데이터세트에서 사전학습한 경우 평가된 기준선들 중 SLAKE에서 세 번째에 위치한다.
- MedTrinity-25M에서의 사전학습은 다운스트림 VQA 벤치마크에서 약 10.75%의 개선(VQA-RAD), 6.1%의 개선(SLAKE), 13.25%의 개선(PathVQA)을 가져왔다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.