Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

Jungo Kasai, Yuhei Kasai|arXiv (Cornell University)|2023. 03. 31.
Artificial Intelligence in Healthcare and Education인용 수 50
한 줄 요약

GPT-4는 Japan’s national medical licensing exams를 합격하여 ChatGPT 및 GPT-3를 능가하지만 여전히 다수의 의대생보다 뒤처리고, 일본어 토큰화로 인한 금지된 선택지와 더 높은 비용 같은 한계가 나타난다.

ABSTRACT

As large language models (LLMs) gain popularity among speakers of diverse languages, we believe that it is crucial to benchmark them to better understand model behaviors, failures, and limitations in languages beyond English. In this work, we evaluate LLM APIs (ChatGPT, GPT-3, and GPT-4) on the Japanese national medical licensing examinations from the past five years, including the current year. Our team comprises native Japanese-speaking NLP researchers and a practicing cardiologist based in Japan. Our experiments show that GPT-4 outperforms ChatGPT and GPT-3 and passes all six years of the exams, highlighting LLMs' potential in a language that is typologically distant from English. However, our evaluation also exposes critical limitations of the current LLM APIs. First, LLMs sometimes select prohibited choices that should be strictly avoided in medical practice in Japan, such as suggesting euthanasia. Further, our analysis shows that the API costs are generally higher and the maximum context size is smaller for Japanese because of the way non-Latin scripts are currently tokenized in the pipeline. We release our benchmark as Igaku QA as well as all model outputs and exam metadata. We hope that our results and benchmark will spur progress on more diverse applications of LLMs. Our benchmark is available at https://github.com/jungokasai/IgakuQA.

연구 동기 및 목표

  • 대형 언어 모델이 2018–2023년 사이의 일본 national medical practitioners qualifying examination(NMPQE)을 처리할 수 있는지 평가한다.
  • GPT-3, ChatGPT, 그리고 GPT-4의 비영어권, 국가별 의학 시험에서의 성능을 비교한다.
  • 일본 의료 맥락에서 현재 LLM API의 도메인 및 언어 특화 한계를 파악한다.
  • 일본어 원주에서도 Igaku QA를 공개 벤치마크로 출시하여 비영어권 LLM 평가와 개발을 촉진한다.

제안 방법

  • IGaku QA에서 LLM API(GPT-3, ChatGPT, GPT-4)를 폐쇄형 프롬프트로 벤치마크한다.
  • 모델 응답을 유도하기 위해 2006년 일본의 의사 면허 시험에서 세 가지 인-컨텍스트 예시를 사용한다.
  • MCQ 형식과 고정 수치 질문으로 인해 정답 매칭으로 답변을 평가한다.
  • 문제를 영어로 번역하는 ChatGPT-EN 변형을 포함시켜 다국어 프롭핑 효과를 시험한다.
  • 일본어에서 금지된 선택지(禁忌肢), 토큰화 관련 비용, 맥락 창(window) 제한과 같은 요인을 분석한다.
Figure 1: Example problem from the Japanese medical licensing exam where ChatGPT chooses a prohibited choice (禁忌肢) because euthanasia is illegal in Japan . Test takers who choose four or more prohibited choices would fail regardless of their exam total scores (§ 2.2 ). The problem and the ChatGPT ou
Figure 1: Example problem from the Japanese medical licensing exam where ChatGPT chooses a prohibited choice (禁忌肢) because euthanasia is illegal in Japan . Test takers who choose four or more prohibited choices would fail regardless of their exam total scores (§ 2.2 ). The problem and the ChatGPT ou

실험 결과

연구 질문

  • RQ1GPT-4가 2018–2023년 사이의 일본 NMPQE를 합격할 수 있는가?
  • RQ2GPT-3, ChatGPT, GPT-4는 서로 비교했을 때와 Igaku QA에서 의대생의 성과와 비교했을 때 어떤 차이가 있는가?
  • RQ3일본 의료 면허 문제에 현재 LLM이 적용될 때의 주요 한계(예: 금지된 선택지, 지리적/시계 맥락, 토큰화 비용)는 무엇인가?

주요 결과

  • GPT-4가 테스트에 사용된 모델들 중에서 가장 높은 성능을 달성한다.
  • GPT-4는 2018–2023년의 일본 의학 면허 시험 모든 6개 연도에서 합격한다.
  • GPT-4는 의대생의 다수 표결 성능에 비해 상당히 뒤처진다.
  • 일부 LLM 출력은 일본 의학 실무에서 피해야 할 금지된 선택지(禁忌肢)를 가끔 선택한다.
  • 일본어 토큰화는 API 비용을 증가시키고 영어에 비해 최대 맥락 창을 축소한다.
  • ChatGPT-EN(영어 프롬프트/번역)은 일반 ChatGPT보다 종종 우수하여 명시적 번역이 없는 다국어 프롭핑의 한계를 시사한다.
Figure 2: Standard timeline comparisons of the medical licensing processes in Japan and the United States. The Japanese system involves only one national licensing examination (NMPQE) towards the end of the six-year medical school education. The United States Medical Licensing Examination (USMLE) co
Figure 2: Standard timeline comparisons of the medical licensing processes in Japan and the United States. The Japanese system involves only one national licensing examination (NMPQE) towards the end of the six-year medical school education. The United States Medical Licensing Examination (USMLE) co

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.