[论文解读] ChatGPT-3.5, ChatGPT-4, Google Bard, and Microsoft Bing to Improve Health Literacy and Communication in Pediatric Populations and Beyond
本文评估四种大型语言模型(ChatGPT-3.5/4、Google Bard、Microsoft Bing)如何为儿童群体定制健康信息,显示出不同的阅读难度输出和提示行为。
Purpose: Enhanced health literacy has been linked to better health outcomes; however, few interventions have been studied. We investigate whether large language models (LLMs) can serve as a medium to improve health literacy in children and other populations. Methods: We ran 288 conditions using 26 different prompts through ChatGPT-3.5, Microsoft Bing, and Google Bard. Given constraints imposed by rate limits, we tested a subset of 150 conditions through ChatGPT-4. The primary outcome measurements were the reading grade level (RGL) and word counts of output. Results: Across all models, output for basic prompts such as "Explain" and "What is (are)" were at, or exceeded, a 10th-grade RGL. When prompts were specified to explain conditions from the 1st to 12th RGL, we found that LLMs had varying abilities to tailor responses based on RGL. ChatGPT-3.5 provided responses that ranged from the 7th-grade to college freshmen RGL while ChatGPT-4 outputted responses from the 6th-grade to the college-senior RGL. Microsoft Bing provided responses from the 9th to 11th RGL while Google Bard provided responses from the 7th to 10th RGL. Discussion: ChatGPT-3.5 and ChatGPT-4 did better in achieving lower-grade level outputs. Meanwhile Bard and Bing tended to consistently produce an RGL that is at the high school level regardless of prompt. Additionally, Bard's hesitancy in providing certain outputs indicates a cautious approach towards health information. LLMs demonstrate promise in enhancing health communication, but future research should verify the accuracy and effectiveness of such tools in this context. Implications: LLMs face challenges in crafting outputs below a sixth-grade reading level. However, their capability to modify outputs above this threshold provides a potential mechanism to improve health literacy and communication in a pediatric population and beyond.
研究动机与目标
- 研究大型语言模型是否可以成为提升儿童及其他人群健康素养的媒介。
- 量化不同模型和提示下输出(阅读水平与长度)的变化。
- 评估使用大型语言模型将健康信息定制到不同阅读水平的可行性及局限性。
提出的方法
- 在 ChatGPT-3.5、Microsoft Bing 和 Google Bard 上,使用 26 条提示对 288 种条件进行了测试。
- 由于速率限制,使用 ChatGPT-4 对其中的 150 种条件进行了子集测试。
- 主要结果是所生成输出的阅读等级(RGL)和字数。
实验结果
研究问题
- RQ1大型语言模型是否能够有效地将健康信息定制到特定的阅读等级?
- RQ2在不同提示下,不同的 LLM 如何在产生较低与较高阅读水平的输出方面表现?
- RQ3将 LLM 用于儿童健康沟通的局限性有哪些(如准确性、犹豫性等)?
主要发现
- 在所有模型中,像 Explain 或 What is 这类基础提示会产生达到或高于 10th-grade RGL 的输出。
- ChatGPT-3.5 的输出阅读等级范围为从 7th-grade 到 college freshmen。
- ChatGPT-4 的输出阅读等级范围为从 6th-grade 到 college senior RGL。
- Microsoft Bing 的输出阅读等级范围为从 9th 到 11th-grade RGL。
- Google Bard 的输出阅读等级范围为从 7th 到 10th-grade RGL。
- Bard 在提供某些输出时表现出犹豫,显示出对健康信息的谨慎态度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。