[논문 리뷰] Evaluating multiple large language models in pediatric ophthalmology
이 연구는 소아 안과 분야에서 100문항의 다중 선택형 시험을 통해 대규모 언어 모델(GPT-4, ChatGPT(GPT-3.5), PaLM2)의 성능을 의예과 학부생, 수련의, 임상 전문의와 비교하였다. GPT-4는 정확도에서 임상 전문의와 유사했으며, 의예과 학부생을 능가했고, 다른 모델들보다 더 높은 응답 안정성과 자신감을 보였다.
IMPORTANCE The response effectiveness of different large language models (LLMs) and various individuals, including medical students, graduate students, and practicing physicians, in pediatric ophthalmology consultations, has not been clearly established yet. OBJECTIVE Design a 100-question exam based on pediatric ophthalmology to evaluate the performance of LLMs in highly specialized scenarios and compare them with the performance of medical students and physicians at different levels. DESIGN, SETTING, AND PARTICIPANTS This survey study assessed three LLMs, namely ChatGPT (GPT-3.5), GPT-4, and PaLM2, were assessed alongside three human cohorts: medical students, postgraduate students, and attending physicians, in their ability to answer questions related to pediatric ophthalmology. It was conducted by administering questionnaires in the form of test papers through the LLM network interface, with the valuable participation of volunteers. MAIN OUTCOMES AND MEASURES Mean scores of LLM and humans on 100 multiple-choice questions, as well as the answer stability, correlation, and response confidence of each LLM. RESULTS GPT-4 performed comparably to attending physicians, while ChatGPT (GPT-3.5) and PaLM2 outperformed medical students but slightly trailed behind postgraduate students. Furthermore, GPT-4 exhibited greater stability and confidence when responding to inquiries compared to ChatGPT (GPT-3.5) and PaLM2. CONCLUSIONS AND RELEVANCE Our results underscore the potential for LLMs to provide medical assistance in pediatric ophthalmology and suggest significant capacity to guide the education of medical students.
연구 동기 및 목표
- 대규모 언어 모델(Large Language Models, LLMs)의 소아 안과 분야에서 임상 추론 정확도를 평가하기 위해.
- LLM 성능을 의예과 학부생, 수련의, 임상 전문의와 비교하기 위해.
- LLM과 인간 집단 간의 응답 안정성, 자신감 수준, 상관관계 평가하기 위해.
- 특화된 의료 분야에서 LLM의 교육적 및 임상 보조 잠재력 탐색하기 위해.
제안 방법
- 소아 안과 분야의 내용을 바탕으로 100문항의 다중 선택형 시험을 개발하였다.
- 네트워크 인터페이스를 통해 ChatGPT(GPT-3.5), GPT-4, PaLM2의 세 개의 LLM에 시험 문제를 제출하여 응답을 수집하였다.
- 의예과 학부생, 수련의, 임상 전문의의 세 개의 인간 집단이 동일한 시험을 수행하였다.
- 성능는 평균 점수, 응답 안정성, 자신감 수준, 인간 응답과의 상관관계로 측정되었다.
- 통계 분석을 통해 LLM과 인간 집단 간의 주요 지표에서의 유의미한 차이를 비교하였다.
실험 결과
연구 질문
- RQ1다양한 LLM은 소아 안과 분야에서 인간 전문의와 비교해 어떻게 성능을 보였는가?
- RQ2LLM은 의예과 학부생 및 전문의의 진단 정확도에 도달하거나 초월할 수 있는가?
- RQ3LLM의 응답은 인간의 응답과 비교해 얼마나 안정적이고 자신감 있는가?
- RQ4LLM과 인간의 임상 추론 과제 수행 능력 간 상관관계는 어떠한가?
주요 결과
- GPT-4는 임상 전문의와 유사한 평균 점수를 기록하여 소아 안과 분야에서 높은 임상 추론 정확도를 보였다.
- ChatGPT(GPT-3.5)와 PaLM2는 의예과 학부생을 능가했지만, 수련의보다 略적으로 낮은 점수를 기록했다.
- GPT-4는 ChatGPT(GPT-3.5)와 PaLM2보다 더 높은 응답 안정성과 자신감 점수를 보였다.
- GPT-4와 임상 전문의 간 성과 상관관계는 다른 LLM-인간 쌍보다 더 강했다.
- LLM, 특히 GPT-4는 100문항의 시험 세트 전반에서 일관되고 신뢰할 수 있는 성능를 보였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.