Skip to main content
QUICK REVIEW

[논문 리뷰] ChatGPT-3.5, ChatGPT-4, Google Bard, and Microsoft Bing to Improve Health Literacy and Communication in Pediatric Populations and Beyond

Kanhai Amin, Linda C. Mayes|arXiv (Cornell University)|2023. 11. 16.
Health Literacy and Information Accessibility인용 수 10
한 줄 요약

요약: 본 논문은 네 가지 LLM(ChatGPT-3.5/4, Google Bard, Microsoft Bing)이 소아 인구를 대상으로 건강 정보를 얼마나 맞춤화하는지 평가하고, 읽기 등급 출력과 프롬프트 동작의 차이를 보여준다.

ABSTRACT

Purpose: Enhanced health literacy has been linked to better health outcomes; however, few interventions have been studied. We investigate whether large language models (LLMs) can serve as a medium to improve health literacy in children and other populations. Methods: We ran 288 conditions using 26 different prompts through ChatGPT-3.5, Microsoft Bing, and Google Bard. Given constraints imposed by rate limits, we tested a subset of 150 conditions through ChatGPT-4. The primary outcome measurements were the reading grade level (RGL) and word counts of output. Results: Across all models, output for basic prompts such as "Explain" and "What is (are)" were at, or exceeded, a 10th-grade RGL. When prompts were specified to explain conditions from the 1st to 12th RGL, we found that LLMs had varying abilities to tailor responses based on RGL. ChatGPT-3.5 provided responses that ranged from the 7th-grade to college freshmen RGL while ChatGPT-4 outputted responses from the 6th-grade to the college-senior RGL. Microsoft Bing provided responses from the 9th to 11th RGL while Google Bard provided responses from the 7th to 10th RGL. Discussion: ChatGPT-3.5 and ChatGPT-4 did better in achieving lower-grade level outputs. Meanwhile Bard and Bing tended to consistently produce an RGL that is at the high school level regardless of prompt. Additionally, Bard's hesitancy in providing certain outputs indicates a cautious approach towards health information. LLMs demonstrate promise in enhancing health communication, but future research should verify the accuracy and effectiveness of such tools in this context. Implications: LLMs face challenges in crafting outputs below a sixth-grade reading level. However, their capability to modify outputs above this threshold provides a potential mechanism to improve health literacy and communication in a pediatric population and beyond.

연구 동기 및 목표

  • 대형 언어 모델이 아동 및 기타 인구의 건강 literacy를 개선하는 매개체가 될 수 있는지 조사한다.
  • 모듈 간 출력물의 읽기 수준 및 길이가 프롬프트에 따라 어떻게 달라지는지 정량화한다.
  • 다양한 읽기 수준에 맞춘 건강 정보를 LLM으로 맞춤화하는 가능성과 한계를 평가한다.

제안 방법

  • ChatGPT-3.5, Microsoft Bing, Google Bard 간 26개의 프롬프트로 288 조건을 수행했다.
  • Rate limit으로 인해 ChatGPT-4에서 150개의 조건의 하위 집합을 테스트했다.
  • 주요 결과 측정 항목은 읽기 등급(RGL)과 생성 출력의 단어 수였다.

실험 결과

연구 질문

  • RQ1LLMs가 특정 읽기 등급으로 건강 정보를 효과적으로 맞춤화할 수 있는가?
  • RQ2다양한 LLM이 프롬프트별로 낮은 읽기 수준 출력 vs 높은 읽기 수준 출력을 내는 데 어떤 차이가 있는가?
  • RQ3소아 건강 커뮤니케이션에 LLM을 사용함에 있어 정확성, 망설임 등 한계는 무엇인가?

주요 결과

  • 모델 전반에서 Explain 또는 What is 같은 기본 프롬프트는 10th-grade RGL에 도달하거나 그 이상으로의 출력을 유도한다.
  • ChatGPT-3.5는 7th-grade에서 college freshmen RGL까지의 출력을 생성했다.
  • ChatGPT-4는 6th-grade에서 college senior RGL까지의 출력을 생성했다.
  • Microsoft Bing은 9th에서 11th-grade RGL까지의 출력을 생성했다.
  • Google Bard는 7th에서 10th-grade RGL까지의 출력을 생성했다.
  • Bard는 특정 출력 제공에 있어 망설임을 보였으며, 건강 정보에 대해 신중한 접근을 나타냈다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.