Skip to main content
QUICK REVIEW

[Paper Review] ChatGPT-3.5, ChatGPT-4, Google Bard, and Microsoft Bing to Improve Health Literacy and Communication in Pediatric Populations and Beyond

Kanhai Amin, Linda C. Mayes|arXiv (Cornell University)|Nov 16, 2023
Health Literacy and Information Accessibility10 citations
TL;DR

The paper evaluates how four LLMs (ChatGPT-3.5/4, Google Bard, Microsoft Bing) tailor health information for pediatric populations, showing varying reading level outputs and prompting behaviors.

ABSTRACT

Purpose: Enhanced health literacy has been linked to better health outcomes; however, few interventions have been studied. We investigate whether large language models (LLMs) can serve as a medium to improve health literacy in children and other populations. Methods: We ran 288 conditions using 26 different prompts through ChatGPT-3.5, Microsoft Bing, and Google Bard. Given constraints imposed by rate limits, we tested a subset of 150 conditions through ChatGPT-4. The primary outcome measurements were the reading grade level (RGL) and word counts of output. Results: Across all models, output for basic prompts such as "Explain" and "What is (are)" were at, or exceeded, a 10th-grade RGL. When prompts were specified to explain conditions from the 1st to 12th RGL, we found that LLMs had varying abilities to tailor responses based on RGL. ChatGPT-3.5 provided responses that ranged from the 7th-grade to college freshmen RGL while ChatGPT-4 outputted responses from the 6th-grade to the college-senior RGL. Microsoft Bing provided responses from the 9th to 11th RGL while Google Bard provided responses from the 7th to 10th RGL. Discussion: ChatGPT-3.5 and ChatGPT-4 did better in achieving lower-grade level outputs. Meanwhile Bard and Bing tended to consistently produce an RGL that is at the high school level regardless of prompt. Additionally, Bard's hesitancy in providing certain outputs indicates a cautious approach towards health information. LLMs demonstrate promise in enhancing health communication, but future research should verify the accuracy and effectiveness of such tools in this context. Implications: LLMs face challenges in crafting outputs below a sixth-grade reading level. However, their capability to modify outputs above this threshold provides a potential mechanism to improve health literacy and communication in a pediatric population and beyond.

Motivation & Objective

  • Investigate whether large language models can serve as a medium to improve health literacy in children and other populations.
  • Quantify how outputs (reading level and length) vary across models and prompts.
  • Assess the feasibility and limitations of using LLMs to tailor health information to different reading levels.

Proposed method

  • Ran 288 conditions using 26 prompts across ChatGPT-3.5, Microsoft Bing, and Google Bard.
  • Tested a subset of 150 conditions with ChatGPT-4 due to rate limits.
  • Primary outcomes measured were reading grade level (RGL) and word counts of generated output.

Experimental results

Research questions

  • RQ1Can LLMs tailor health information to specific reading grade levels effectively?
  • RQ2How do different LLMs perform in producing lower vs higher reading level outputs across prompts?
  • RQ3What are the limitations (e.g., accuracy, hesitancy) of using LLMs for pediatric health communication?

Key findings

  • Across models, basic prompts like Explain or What is lead to outputs at or above a 10th-grade RGL.
  • ChatGPT-3.5 produced outputs ranging from 7th-grade to college freshmen RGL.
  • ChatGPT-4 produced outputs ranging from 6th-grade to college senior RGL.
  • Microsoft Bing produced outputs ranging from 9th to 11th-grade RGL.
  • Google Bard produced outputs ranging from 7th to 10th-grade RGL.
  • Bard showed hesitancy in providing certain outputs, indicating a cautious approach to health information.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.