[논문 리뷰] Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology
본 연구는 자동화된 불임 병력 수집을 위해 ChatGPT-4o와 ChatGPT-4o-mini를 평가하며, 4o-mini가 추출 완전성에서 우수하고 다른 지표의 차이는 작다.
Effective physician-patient communications in pre-diagnostic environments, and most specifically in complex and sensitive medical areas such as infertility, are critical but consume a lot of time and, therefore, cause clinic workflows to become inefficient. Recent advancements in Large Language Models (LLMs) offer a potential solution for automating conversational medical history-taking and improving diagnostic accuracy. This study evaluates the feasibility and performance of LLMs in those tasks for infertility cases. An AI-driven conversational system was developed to simulate physician-patient interactions with ChatGPT-4o and ChatGPT-4o-mini. A total of 70 real-world infertility cases were processed, generating 420 diagnostic histories. Model performance was assessed using F1 score, Differential Diagnosis (DDs) Accuracy, and Accuracy of Infertility Type Judgment (ITJ). ChatGPT-4o-mini outperformed ChatGPT-4o in information extraction accuracy (F1 score: 0.9258 vs. 0.9029, p = 0.045, d = 0.244) and demonstrated higher completeness in medical history-taking (97.58% vs. 77.11%), suggesting that ChatGPT-4o-mini is more effective in extracting detailed patient information, which is critical for improving diagnostic accuracy. In contrast, ChatGPT-4o performed slightly better in differential diagnosis accuracy (2.0524 vs. 2.0048, p > 0.05). ITJ accuracy was higher in ChatGPT-4o-mini (0.6476 vs. 0.5905) but with lower consistency (Cronbach's $α$ = 0.562), suggesting variability in classification reliability. Both models demonstrated strong feasibility in automating infertility history-taking, with ChatGPT-4o-mini excelling in completeness and extraction accuracy. In future studies, expert validation for accuracy and dependability in a clinical setting, AI model fine-tuning, and larger datasets with a mix of cases of infertility have to be prioritized.
연구 동기 및 목표
- 산부인과 분야에서 LLM이 불임 의료 이력 수집을 자동화하는 타당성을 평가한다.
- ChatGPT-4o와 ChatGPT-4o-mini 간 정보 추출 및 진단지원 성능을 비교한다.
- 이력 수집의 완전성과 차별 진단(DDs) 및 불임 유형 판단(DTJ) 신뢰성을 평가한다.
제안 방법
- 의사-환자 상호작용을 시뮬레이션하기 위한 AI 기반 대화 시스템을 개발한다.
- 실제 불임 사례 70건을 처리하여 420개의 진단 병력을 생성한다.
- 정보 추출의 F1 점수, 차별 진단(DDs) 정확도, 불임 유형 판단(ITJ) 정확도로 성능을 평가한다.
- 추출, 완전성 및 진단 지표에서 ChatGPT-4o와 ChatGPT-4o-mini를 비교한다.
실험 결과
연구 질문
- RQ1LLM 기반 시스템이 자동으로 정확하고 완전한 불임 의료 이력을 생성할 수 있는가?
- RQ2불임 사례에서 정보 추출, DDs 정확도 및 ITJ 정확도 측면에서 ChatGPT-4o와 ChatGPT-4o-mini는 어떻게 비교되는가?
주요 결과
- ChatGPT-4o-mini는 ChatGPT-4o( F1 0.9029 )보다 정보 추출 정확도가 더 높다(F1 0.9258), p = 0.045, d = 0.244.
- ChatGPT-4o-mini는 의료 이력 수집의 완전성이 더 높음을 보여준다(97.58%) 대 ChatGPT-4o(77.11%).
- ChatGPT-4o가 차별 진단 정확도에서 약간 더 우수한 것으로 보이며(2.0524) ChatGPT-4o-mini(2.0048)보다, p > 0.05로 나타난다.
- ITJ 정확도는 ChatGPT-4o-mini(0.6476)가 ChatGPT-4o(0.5905)보다 높지만 일관성은 더 낮다(Cronbach’s α = 0.562).
- 두 모델 모두 불임 이력 수집 자동화에 대한 강한 타당성을 보여준다; 4o-mini가 완전성과 추출에서 우수하며; 임상 검증과 더 큰 데이터 세트가 필요하다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.