Skip to main content
QUICK REVIEW

[Paper Review] Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology

Dandan Liu, Ying Long|ArXiv.org|Mar 31, 2025
Artificial Intelligence in Healthcare and Education3 citations
TL;DR

The study evaluates ChatGPT-4o and ChatGPT-4o-mini for automated infertility history-taking, finding that 4o-mini excels in extraction completeness, with modest differences in other metrics.

ABSTRACT

Effective physician-patient communications in pre-diagnostic environments, and most specifically in complex and sensitive medical areas such as infertility, are critical but consume a lot of time and, therefore, cause clinic workflows to become inefficient. Recent advancements in Large Language Models (LLMs) offer a potential solution for automating conversational medical history-taking and improving diagnostic accuracy. This study evaluates the feasibility and performance of LLMs in those tasks for infertility cases. An AI-driven conversational system was developed to simulate physician-patient interactions with ChatGPT-4o and ChatGPT-4o-mini. A total of 70 real-world infertility cases were processed, generating 420 diagnostic histories. Model performance was assessed using F1 score, Differential Diagnosis (DDs) Accuracy, and Accuracy of Infertility Type Judgment (ITJ). ChatGPT-4o-mini outperformed ChatGPT-4o in information extraction accuracy (F1 score: 0.9258 vs. 0.9029, p = 0.045, d = 0.244) and demonstrated higher completeness in medical history-taking (97.58% vs. 77.11%), suggesting that ChatGPT-4o-mini is more effective in extracting detailed patient information, which is critical for improving diagnostic accuracy. In contrast, ChatGPT-4o performed slightly better in differential diagnosis accuracy (2.0524 vs. 2.0048, p > 0.05). ITJ accuracy was higher in ChatGPT-4o-mini (0.6476 vs. 0.5905) but with lower consistency (Cronbach's $α$ = 0.562), suggesting variability in classification reliability. Both models demonstrated strong feasibility in automating infertility history-taking, with ChatGPT-4o-mini excelling in completeness and extraction accuracy. In future studies, expert validation for accuracy and dependability in a clinical setting, AI model fine-tuning, and larger datasets with a mix of cases of infertility have to be prioritized.

Motivation & Objective

  • Assess feasibility of LLMs to automate infertility medical history-taking in obstetrics/gynecology.
  • Compare information extraction and diagnostic-support performance between ChatGPT-4o and ChatGPT-4o-mini.
  • Evaluate completeness of history-taking and reliability of differential diagnosis and infertility-type judgments.

Proposed method

  • Develop AI-driven conversational system to simulate physician-patient interactions.
  • Process 70 real-world infertility cases to generate 420 diagnostic histories.
  • Assess performance using F1 score for information extraction, Differential Diagnosis (DDs) Accuracy, and Infertility Type Judgment (ITJ) Accuracy.
  • Compare ChatGPT-4o and ChatGPT-4o-mini on extraction, completeness, and diagnostic metrics.

Experimental results

Research questions

  • RQ1Can LLM-based systems automatically generate accurate and complete infertility medical histories?
  • RQ2How do ChatGPT-4o and ChatGPT-4o-mini compare in information extraction, DDs accuracy, and ITJ accuracy for infertility cases?

Key findings

  • ChatGPT-4o-mini has higher information extraction accuracy (F1 0.9258) than ChatGPT-4o (F1 0.9029), p = 0.045, d = 0.244.
  • ChatGPT-4o-mini demonstrates higher completeness in medical history-taking (97.58%) versus ChatGPT-4o (77.11%).
  • ChatGPT-4o shows slightly better differential diagnosis accuracy (2.0524) than ChatGPT-4o-mini (2.0048), with p > 0.05.
  • ITJ accuracy is higher for ChatGPT-4o-mini (0.6476) than for ChatGPT-4o (0.5905), but consistency is lower (Cronbach’s α = 0.562).
  • Both models show strong feasibility for automating infertility history-taking; 4o-mini excels in completeness and extraction; clinical validation and larger datasets are needed.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.