[Paper Review] Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology
The study evaluates ChatGPT-4o and ChatGPT-4o-mini for automated infertility history-taking, finding that 4o-mini excels in extraction completeness, with modest differences in other metrics.
Effective physician-patient communications in pre-diagnostic environments, and most specifically in complex and sensitive medical areas such as infertility, are critical but consume a lot of time and, therefore, cause clinic workflows to become inefficient. Recent advancements in Large Language Models (LLMs) offer a potential solution for automating conversational medical history-taking and improving diagnostic accuracy. This study evaluates the feasibility and performance of LLMs in those tasks for infertility cases. An AI-driven conversational system was developed to simulate physician-patient interactions with ChatGPT-4o and ChatGPT-4o-mini. A total of 70 real-world infertility cases were processed, generating 420 diagnostic histories. Model performance was assessed using F1 score, Differential Diagnosis (DDs) Accuracy, and Accuracy of Infertility Type Judgment (ITJ). ChatGPT-4o-mini outperformed ChatGPT-4o in information extraction accuracy (F1 score: 0.9258 vs. 0.9029, p = 0.045, d = 0.244) and demonstrated higher completeness in medical history-taking (97.58% vs. 77.11%), suggesting that ChatGPT-4o-mini is more effective in extracting detailed patient information, which is critical for improving diagnostic accuracy. In contrast, ChatGPT-4o performed slightly better in differential diagnosis accuracy (2.0524 vs. 2.0048, p > 0.05). ITJ accuracy was higher in ChatGPT-4o-mini (0.6476 vs. 0.5905) but with lower consistency (Cronbach's $α$ = 0.562), suggesting variability in classification reliability. Both models demonstrated strong feasibility in automating infertility history-taking, with ChatGPT-4o-mini excelling in completeness and extraction accuracy. In future studies, expert validation for accuracy and dependability in a clinical setting, AI model fine-tuning, and larger datasets with a mix of cases of infertility have to be prioritized.
Motivation & Objective
- Assess feasibility of LLMs to automate infertility medical history-taking in obstetrics/gynecology.
- Compare information extraction and diagnostic-support performance between ChatGPT-4o and ChatGPT-4o-mini.
- Evaluate completeness of history-taking and reliability of differential diagnosis and infertility-type judgments.
Proposed method
- Develop AI-driven conversational system to simulate physician-patient interactions.
- Process 70 real-world infertility cases to generate 420 diagnostic histories.
- Assess performance using F1 score for information extraction, Differential Diagnosis (DDs) Accuracy, and Infertility Type Judgment (ITJ) Accuracy.
- Compare ChatGPT-4o and ChatGPT-4o-mini on extraction, completeness, and diagnostic metrics.
Experimental results
Research questions
- RQ1Can LLM-based systems automatically generate accurate and complete infertility medical histories?
- RQ2How do ChatGPT-4o and ChatGPT-4o-mini compare in information extraction, DDs accuracy, and ITJ accuracy for infertility cases?
Key findings
- ChatGPT-4o-mini has higher information extraction accuracy (F1 0.9258) than ChatGPT-4o (F1 0.9029), p = 0.045, d = 0.244.
- ChatGPT-4o-mini demonstrates higher completeness in medical history-taking (97.58%) versus ChatGPT-4o (77.11%).
- ChatGPT-4o shows slightly better differential diagnosis accuracy (2.0524) than ChatGPT-4o-mini (2.0048), with p > 0.05.
- ITJ accuracy is higher for ChatGPT-4o-mini (0.6476) than for ChatGPT-4o (0.5905), but consistency is lower (Cronbach’s α = 0.562).
- Both models show strong feasibility for automating infertility history-taking; 4o-mini excels in completeness and extraction; clinical validation and larger datasets are needed.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.