[Paper Review] A Study of Generative Large Language Model for Medical Research and Healthcare
The paper develops a clinical generative LLM, GatorTronGPT, trained on 277B words with 20B parameters, and shows synthetic NLP models outperforming real-world clinical-text models while physicians cannot distinguish AI from humans in a Turing test.
There is enormous enthusiasm and concerns in using large language models (LLMs) in healthcare, yet current assumptions are all based on general-purpose LLMs such as ChatGPT. This study develops a clinical generative LLM, GatorTronGPT, using 277 billion words of mixed clinical and English text with a GPT-3 architecture of 20 billion parameters. GatorTronGPT improves biomedical natural language processing for medical research. Synthetic NLP models trained using GatorTronGPT generated text outperform NLP models trained using real-world clinical text. Physicians Turing test using 1 (worst) to 9 (best) scale shows that there is no significant difference in linguistic readability (p = 0.22; 6.57 of GatorTronGPT compared with 6.93 of human) and clinical relevance (p = 0.91; 7.0 of GatorTronGPT compared with 6.97 of human) and that physicians cannot differentiate them (p < 0.001). This study provides insights on the opportunities and challenges of LLMs for medical research and healthcare.
Motivation & Objective
- Motivate the use of large language models in medical research and healthcare beyond general-purpose LLMs.
- Develop a clinical generative LLM (GatorTronGPT) tailored to medical text using large-scale mixed clinical and English data.
- Evaluate the performance of GatorTronGPT on biomedical NLP tasks and compare synthetic vs real-world clinical-text models.
- Assess physician perception of AI-generated medical text through a Turing-test style evaluation.
Proposed method
- Construct GatorTronGPT with a GPT-3 architecture comprising 20 billion parameters.
- Train on a corpus of 277 billion words of mixed clinical and English text.
- Evaluate NLP performance on biomedical tasks and compare against models trained on real clinical text.
- Generate synthetic NLP models using text from GatorTronGPT to benchmark against models trained on real clinical data.
- Conduct a physician Turing-test style evaluation for linguistic readability and clinical relevance, using a 1–9 scale.
Experimental results
Research questions
- RQ1Can a clinical generative LLM trained on mixed clinical and English data outperform models trained only on real-world clinical text in biomedical NLP tasks?
- RQ2Are AI-generated medical text outputs indistinguishable from human-authored text in readability and clinical relevance to physicians?
- RQ3What are the opportunities and challenges of deploying LLMs in medical research and healthcare based on empirical evaluations?
Key findings
- Synthetic NLP models trained on GatorTronGPT-generated text outperform NLP models trained on real-world clinical text.
- Physician Turing-test results show no significant difference in linguistic readability between GatorTronGPT (6.57) and human (6.93) samples (p = 0.22).
- Physician Turing-test results show no significant difference in clinical relevance between GatorTronGPT (7.0) and human (6.97) samples (p = 0.91).
- Physicians cannot reliably distinguish AI-generated from human-authored outputs (p < 0.001).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.