Skip to main content
QUICK REVIEW

[논문 리뷰] A Foundational Multimodal Vision Language AI Assistant for Human Pathology

Ming Y. Lu, Bowen Chen|arXiv (Cornell University)|2023. 12. 13.
Artificial Intelligence in Healthcare and Education인용 수 15
한 줄 요약

PathChat은 UNI에서 파생된 비전 인코더와 13B LLM에 연결된 병리학용 비전-언어 AI 어시스턴트이며, 257k 병리 지시사항으로 학습되었고, 임상 맥락이 제공될 때 특히 다지선다형 및 개방형 병리 문제에서 기본 모델보다 우수하다.

ABSTRACT

The field of computational pathology has witnessed remarkable progress in the development of both task-specific predictive models and task-agnostic self-supervised vision encoders. However, despite the explosive growth of generative artificial intelligence (AI), there has been limited study on building general purpose, multimodal AI assistants tailored to pathology. Here we present PathChat, a vision-language generalist AI assistant for human pathology using an in-house developed foundational vision encoder pretrained on 100 million histology images from over 100,000 patient cases and 1.18 million pathology image-caption pairs. The vision encoder is then combined with a pretrained large language model and the whole system is finetuned on over 250,000 diverse disease agnostic visual language instructions. We compare PathChat against several multimodal vision language AI assistants as well as GPT4V, which powers the commercially available multimodal general purpose AI assistant ChatGPT-4. When relevant clinical context is provided with the histology image, PathChat achieved a diagnostic accuracy of 87% on multiple-choice questions based on publicly available cases of diverse tissue origins and disease models. Additionally, using open-ended questions and human expert evaluation, we found that overall PathChat produced more accurate and pathologist-preferable responses to diverse queries related to pathology. As an interactive and general vision language AI assistant that can flexibly handle both visual and natural language inputs, PathChat can potentially find impactful applications in pathology education, research, and human-in-the-loop clinical decision making.

연구 동기 및 목표

  • 병리학에 맞춘 범용 다중모달 AI 어시스턴트를 동기 부여하고 활용 가능하게 한다.
  • 병리학 기초 비전 인코더를 대형 언어 모델과 결합하여 PathChat을 개발한다.
  • 미세조정을 위한 대규모 병리학 중심 지시 데이터셋을 큐레이션하고 활용한다.
  • 진단 및 상호작용 과제 전반에서 PathChat을 오픈 소스 및 상용 다중모달 모델과 평가한다.

제안 방법

  • 100백만 개가 넘는 조직학 이미지로 사전 학습된 시작 비전 인코더로 UNI를 사용한다.
  • 1.18백만 개의 병리학 이미지-캡션 쌍에서 비전-언어 정렬 인코더(CONCH-Large)를 미세조정한다.
  • 멀티모달 프로젝터 모듈을 통해 비전 인코더를 13B 매개변수의 사전학습 LLM에 연결한다.
  • 합쳐진 MLLM을 257k개의 병리학 지시문 데이터셋(PathChatInstruct)에서 미세조정한다.
  • PathChat을 다중선택 진단 문제와 개방형 질문에서 LLaVA, LLaVA-Med 및 GPT4V와 비교 평가하되 맥락 인식 시나리오를 포함한다.

실험 결과

연구 질문

  • RQ1PathChatInstruct를 넘어선 작업별 미세조정 없이 제로샷 또는 파샷 설정에서 PathChat이 조직병리 이미지를 진단할 수 있는가?
  • RQ2현미경 기반 진단 및 개방형 병리 질문에서 PathChat이 일반 목적 및 생물의학적으로 특화된 MLLMs에 비해 어떻게 성능을 보이는가?
  • RQ3임상 맥락을 제공하면 PathChat 어시스턴트의 진단 정확도와 활용도 가 향상되는가?
  • RQ4현미경, 진단, 임상 지식, 보조 검사 범주에서 PathChat의 비교 강점과 약점은 무엇인가?

주요 결과

  • PathChat은 이미지만 있는 다지선다형에서 70.8%의 정확도, 임상 맥락을 제공하면 결합된 병리 벤치마크에서 81.2%를 달성한다.
  • PathChat은 이미지 전용 및 이미지-맥락 설정 모두에서 LLaVA 1.5 및 LLaVA-Med를 능가한다.
  • 개방형 질문에서 PathChat은 전체 정확도 86.1%를 달성하여 GPT4V(59.1%), LLaVA 1.5(42.6%), LLaVA-Med(50.4%)를 능가한다.
  • PathChat은 현미경 및 진단 범주에서 특히 강력한 성능을 보이며 해당 영역에서 GPT4V보다 높은 정확도를 보이고, 반면 GPT4V는 임상 및 보조 검사 질문에서 뛰어나다.
  • PathChat은 대화형의 다회전 상호작용과 휴먼-인-루프 차등 진단 워크플로우를 지원한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.