Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis

Dimitrios P. Panagoulias, Maria Virvou|arXiv (Cornell University)|2024. 01. 28.
Radiomics and Machine Learning in Medical Imaging인용 수 11
한 줄 요약

이 논문은 다중모달 상호작용과 도메인 특화 분석을 결합한 두 단계 LLM 평가 패러다임을 제안하여 이미지가 포함된 병리학 MCQ에서 GPT-4-Vision-Preview를 평가하고 약 84%의 정확도를 달성하며 NER 및 지식 그래프를 통해 통찰을 추출한다.

ABSTRACT

Large language models (LLMs) constitute a breakthrough state-of-the-art Artificial Intelligence technology which is rapidly evolving and promises to aid in medical diagnosis. However, the correctness and the accuracy of their returns has not yet been properly evaluated. In this work, we propose an LLM evaluation paradigm that incorporates two independent steps of a novel methodology, namely (1) multimodal LLM evaluation via structured interactions and (2) follow-up, domain-specific analysis based on data extracted via the previous interactions. Using this paradigm, (1) we evaluate the correctness and accuracy of LLM-generated medical diagnosis with publicly available multimodal multiple-choice questions(MCQs) in the domain of Pathology and (2) proceed to a systemic and comprehensive analysis of extracted results. We used GPT-4-Vision-Preview as the LLM to respond to complex, medical questions consisting of both images and text, and we explored a wide range of diseases, conditions, chemical compounds, and related entity types that are included in the vast knowledge domain of Pathology. GPT-4-Vision-Preview performed quite well, scoring approximately 84\% of correct diagnoses. Next, we further analyzed the findings of our work, following an analytical approach which included Image Metadata Analysis, Named Entity Recognition and Knowledge Graphs. Weaknesses of GPT-4-Vision-Preview were revealed on specific knowledge paths, leading to a further understanding of its shortcomings in specific areas. Our methodology and findings are not limited to the use of GPT-4-Vision-Preview, but a similar approach can be followed to evaluate the usefulness and accuracy of other LLMs and, thus, improve their use with further optimization.

연구 동기 및 목표

  • 의학 분야에서 다중모달 LLM의 평가를 위한 두 단계 패러다임 제안(다중모달 평가와 도메인 특화 분석).
  • 이미지+텍스트 MCQ를 사용한 LLM 생성 병리 진단의 정확도와 정확성 평가.
  • IMA, NER, 지식 그래프를 활용한 결과 분석으로 약점을 식별하고 파인튜닝을 유도.
  • 공개 병리학 MCQ에서 GPT-4-Vision-Preview를 활용한 방법론 시연.

제안 방법

  • 사전 규칙이 포함된 이미지+텍스트 MCQ를 이용한 구조화된 다중모달 상호작용.
  • 특정 응답 형식과 간결한 답변을 강제하기 위한 프롬프트 엔지니어링.
  • Image Metadata Analysis(IMA), Named Entity Recognition(NER), Knowledge Graphs(KGs)를 포함한 데이터 추출.
  • 올바른 응답과 오답 응답 및 설명으로부터 파인튜닝 요구사항을 도출하는 도메인 특화 분석.
Figure 1: Multimodal LLM evaluation
Figure 1: Multimodal LLM evaluation

실험 결과

연구 질문

  • RQ1의학 이미지와 증상 기반 텍스트를 결합하여 진단할 때 LLM의 정확도는 얼마나 되는가?
  • RQ2IMA, NER, KG 분석을 통해 LLM이 드러내는 약점이나 지식 경로는 무엇인가?
  • RQ3평가 방법론이 도메인 특화 의료 업무를 위한 표적 파인튜닝이나 재훈련을 안내할 수 있는가?
  • RQ4이 접근법이 GPT-4-Vision-Preview를 넘어 다른 LLM에도 일반화될 수 있는가?

주요 결과

  • GPT-4-Vision-Preview가 다중모달 병리학 MCQ에서 약 84%의 올바른 진단을 달성했다.
  • 정답 응답: 66/79; 오답 응답: 13/79 중 79개 이미지-질문 항목에서.
  • 병리학 하위 도메인은 다양한 성능을 보였으며 예: 일반 병리학: 죽상동맥경화증 및 혈전증 8/10, 세포 손상 7/10, 면역병리학 9/10, 염증 6/10, 종양 형성 10/10; 기관 계통 병리: 순환기 9/9, 피부병리 9/10, 내분비 8/10.
  • IMA는 이미지 도메인 약점을 주로 심혈관, 피부 및 내분비 이미지에서 확인했다.
  • NER 및 KG 분석은 파인튜닝에서 타깃으로 삼을 특정 엔티티와 지식 경로의 약점을 밝혀냈다.
Figure 2: Domain specific Analysis
Figure 2: Domain specific Analysis

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.