[논문 리뷰] Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis
이 연구는 GPT-4V의 다중 모드 의료 진단 능력을 17개 신체 시스템과 8가지 영상 모달리티에 걸쳐 평가하며, 모달리티/해부학 인식에서 강점을 보이나 진단, 보고서 작성, 위치 지정 및 다중 이미지 추론에 큰 차이를 드러낸다.
Driven by the large foundation models, the development of artificial intelligence has witnessed tremendous progress lately, leading to a surge of general interest from the public. In this study, we aim to assess the performance of OpenAI's newest model, GPT-4V(ision), specifically in the realm of multimodal medical diagnosis. Our evaluation encompasses 17 human body systems, including Central Nervous System, Head and Neck, Cardiac, Chest, Hematology, Hepatobiliary, Gastrointestinal, Urogenital, Gynecology, Obstetrics, Breast, Musculoskeletal, Spine, Vascular, Oncology, Trauma, Pediatrics, with images taken from 8 modalities used in daily clinic routine, e.g., X-ray, Computed Tomography (CT), Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Digital Subtraction Angiography (DSA), Mammography, Ultrasound, and Pathology. We probe the GPT-4V's ability on multiple clinical tasks with or without patent history provided, including imaging modality and anatomy recognition, disease diagnosis, report generation, disease localisation. Our observation shows that, while GPT-4V demonstrates proficiency in distinguishing between medical image modalities and anatomy, it faces significant challenges in disease diagnosis and generating comprehensive reports. These findings underscore that while large multimodal models have made significant advancements in computer vision and natural language processing, it remains far from being used to effectively support real-world medical applications and clinical decision-making. All images used in this report can be found in https://github.com/chaoyi-wu/GPT-4V_Medical_Evaluation.
연구 동기 및 목표
- GPT-4V가 의료 영상 모달리티와 해부학을 인식하는 능력을 평가한다.
- 다수의 영상 모달리티에 걸쳐 진단, 보고서 생성 및 로컬라이제이션에서 GPT-4V의 성능을 평가한다.
- 환자 병력과 다중 이미지를 입력으로 사용할 때 GPT-4V 출력에 미치는 영향을 검토한다.
- 영상의학 및 병리학에서 GPT-4V의 임상 사용에 대한 한계와 안전성 고려사항을 식별한다.
제안 방법
- Radiopaedia에서 17개 신체 시스템과 8가지 영상 모달리티에 대한 영상의학 사례를 선택한다.
- 온라인 인터페이스를 통해 최대 4장의 2D 이미지를 GPT-4V에 입력하고 진단, 보고서 작성, 로컬라이제이션 등의 작업을 지시한다.
- 정합성 기준으로 Radiopaedia/PathologyOutlines의 참조 주석을 사용하되 표준 임상 형식의 한계를 주석한다.
- 병리 평가를 위한 두 차례 대화를 수행한다(이미지만, 그다음 이미지와 조직 기원) 및 단계적 로컬라이제이션 작업(존재 여부, 경계 상자, IOU).
- 이미지 강도을 클램프하고 정규화하며 전문가 방사선의 지침에 따라 핵심 슬라이스를 선택한다; 다중 이미지 입력 및 교차 모달 입력을 별도로 평가한다.

실험 결과
연구 질문
- RQ1GPT-4V가 의료 이미지에서 영상 모달리티와 해부학적 구조를 올바르게 인식할 수 있는가?
- RQ2GPT-4V가 의료 영상 내 해부학적 구조나 이상 소견을 로컬라이즈할 수 있는가?
- RQ3GPT-4V가 정확하고 임상적으로 의미 있는 영상의학 또는 병리 보고서를 생성할 수 있는가?
- RQ4환자 병력이나 서로 다른 모달리티의 다중 이미지를 통합할 때 GPT-4V의 성능은 어떠한가?
- RQ5현실 세계의 의료 의사 결정에서 GPT-4V를 사용할 때의 한계와 안전성 고려사항은 무엇인가?
주요 결과
- GPT-4V는 많은 케이스에서 영상 모달리티와 해부학을 인식할 수 있다.
- GPT-4V는 정확한 질병 진단 및 포괄적 보고서 생성에 어려움을 겪는다.
- GPT-4V는 구조화된 보고서를 생성할 수 있지만 내용이 종종 올바르지 않다.
- GPT-4V는 이미지의 문자와 마커를 OCR할 수 있지만 주석을 잘못 해석할 수 있다.
- GPT-4V는 의료 기기와 그 위치를 식별할 수 있다.
- GPT-4V는 다중 이미지를 분석하고 라운드 간 맥락을 유지하는 데 어려움이 있다.
- GPT-4V의 예측은 환자 병력에 크게 영향을 받는다.
- GPT-4V는 해부학적 구조나 이상 소견을 신뢰성 있게 로컬라이즈할 수 없으며(낮은 IOU, 높은 분산).
- 성능은 변화하며 출력에서 일관성 부족과 안전성 우려를 보인다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.