Skip to main content
QUICK REVIEW

[논문 리뷰] Performance of ChatGPT on USMLE: Unlocking the Potential of Large Language Models for AI-Assisted Medical Education

Prabin Sharma, Kisan Thapa|arXiv (Cornell University)|2023. 06. 30.
Artificial Intelligence in Healthcare and Education인용 수 16
한 줄 요약

본 연구는 하버드 해부학 데이터와 의사 판정으로 USMLE 스타일 문제에 대해 ChatGPT를 평가하고, Google보다 맥락 지향적 추론을 더 잘 수행하며, 논리 질문에서 58.8%, 윤리 질문에서 60%를 달성했다.

ABSTRACT

Artificial intelligence is gaining traction in more ways than ever before. The popularity of language models and AI-based businesses has soared since ChatGPT was made available to the general public via OpenAI. It is becoming increasingly common for people to use ChatGPT both professionally and personally. Considering the widespread use of ChatGPT and the reliance people place on it, this study determined how reliable ChatGPT can be for answering complex medical and clinical questions. Harvard University gross anatomy along with the United States Medical Licensing Examination (USMLE) questionnaire were used to accomplish the objective. The paper evaluated the obtained results using a 2-way ANOVA and posthoc analysis. Both showed systematic covariation between format and prompt. Furthermore, the physician adjudicators independently rated the outcome's accuracy, concordance, and insight. As a result of the analysis, ChatGPT-generated answers were found to be more context-oriented and represented a better model for deductive reasoning than regular Google search results. Furthermore, ChatGPT obtained 58.8% on logical questions and 60% on ethical questions. This means that the ChatGPT is approaching the passing range for logical questions and has crossed the threshold for ethical questions. The paper believes ChatGPT and other language learning models can be invaluable tools for e-learners; however, the study suggests that there is still room to improve their accuracy. In order to improve ChatGPT's performance in the future, further research is needed to better understand how it can answer different types of questions.

연구 동기 및 목표

  • USMLE 스타일 평가에 관련된 복잡한 의학 및 임상 질문에 대한 ChatGPT의 신뢰성을 평가한다.
  • 의학 질문 형식에서 ChatGPT의 성능을 전통적 검색 방법(구글)과 비교한다.
  • 통계 분석을 사용하여 질문 형식과 프롬프트가 ChatGPT 성능에 미치는 영향을 평가한다.
  • 의사 판정자들을 참여시켜 AI가 생성한 답변의 정확성, 일치도, 통찰력을 평가한다.

제안 방법

  • 하버드 대학교 해부학 내용과 USMLE 스타일 질문을 평가 자료로 사용한다.
  • 형식과 프롬프트에 따른 성능을 분석하기 위해 이원 분산분석(2-way ANOVA)을 적용한다.
  • 형식과 프롬프트 간의 상호작용 효과를 탐구하기 위한 사후 분석을 수행한다.
  • 의사 판정자들이 독립적으로 정확성, 일치도, 통찰력에 대해 AI 출력물을 평가하도록 한다.

실험 결과

연구 질문

  • RQ1ChatGPT가 논리적 도메인과 윤리적 도메인 전반에 걸쳐 USMLE 스타일 질문에 신뢰성 있게 답할 수 있는가?
  • RQ2질문 형식이나 프롬프트 스타일이 ChatGPT의 성능에 체계적으로 영향을 미치는가?
  • RQ3의사 평가자들은 표준 검색 결과와 비교할 때 ChatGPT 응답의 정확성, 일치도, 통찰력을 어떻게 평가하는가?

주요 결과

  • ChatGPT가 생성한 답변은 Google 검색 결과보다 더 맥락 지향적이었다.
  • ChatGPT는 일반 구글 검색 결과보다 추론적 추론에 있어 더 나은 모습을 보였다.
  • 논리 질문에서 58.8%, 윤리 질문에서 60%를 얻었다.
  • ChatGPT는 논리 질문의 합격 범위에 근접하는 것으로 보이며 윤리 질문은 임계점을 넘은 것으로 보인다.
  • 통계 분석은 형식과 프롬프트 간의 체계적 공변성을 시사했으며, 이는 이원 분산분석과 사후 분석에서 나타났다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.