Skip to main content
QUICK REVIEW

[Paper Review] Supplementary Material for "The Unreasonable Effectiveness of Large Language Models in Issuing Just-in-Time Adaptive Interventions: Fostering Physical Activity in a Prospective Cardiac Rehabilitation Setting"

David Haag, Devender Kumar|arXiv (Cornell University)|Feb 12, 2024
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This study evaluates the use of GPT-4 to generate personalized, context-aware Just-in-Time Adaptive Interventions (JITAIs) for promoting physical activity in cardiac rehabilitation. Using 450 patient-specific interventions across diverse personas and contexts, GPT-4 outperformed both laypersons and healthcare professionals in appropriateness, engagement, effectiveness, and professionalism, demonstrating LLMs' potential to revolutionize scalable, adaptive digital health support.

ABSTRACT

This upload contains supplementary materials for the publication: "The Last JITAI? - The Unreasonable Effectiveness of Large Language Models in Issuing Just-in-Time Adaptive Interventions: Fostering Physical Activity in a Prospective Cardiac Rehabilitation Setting" Supplement 1: Persona and context information used as basis for JITAI decision-making and content generation Supplement 2: Instructions given to JITAI generators Supplement 3: Data and analysis skripts (contains a Read_me file specifying further details)

Motivation & Objective

  • To assess the feasibility and quality of Large Language Models (LLMs) in generating Just-in-Time Adaptive Interventions (JITAIs) for physical activity promotion in cardiac rehabilitation.
  • To address scalability and flexibility limitations of traditional JITAI implementation models that rely on manual or rule-based design.
  • To compare LLM-generated JITAI suggestions against those created by laypersons and healthcare professionals in real-world clinical relevance and quality.
  • To evaluate the performance of GPT-4 in generating context-sensitive, patient-tailored behavioral health messages across varying clinical severities and situational contexts.

Proposed method

  • The study used GPT-4 to generate 450 JITAI decisions and messages based on five distinct context sets per patient persona, representing varying levels of cardiovascular disease severity.
  • Three patient personas were defined, each reflecting different clinical profiles and behavioral needs in cardiac rehabilitation.
  • Context sets included real-time factors such as activity level, mood, social context, and clinical status, simulating dynamic, real-world conditions.
  • Generated messages were systematically evaluated against human-generated interventions from 10 laypersons and 10 healthcare professionals using four key metrics: appropriateness, engagement, effectiveness, and professionalism.
  • Evaluation was conducted via expert and user-based assessment, ensuring alignment with clinical guidelines and patient-centered design principles.
  • The LLM was prompted with structured templates and clinical context to ensure consistency and relevance in message generation.

Experimental results

Research questions

  • RQ1Can GPT-4 generate JITAI messages that are as clinically appropriate and engaging as those created by healthcare professionals in a cardiac rehabilitation context?
  • RQ2How does the quality of LLM-generated JITAIs compare to those produced by laypersons in terms of personalization and clinical relevance?
  • RQ3To what extent can LLMs scale personalized, context-aware interventions across diverse patient profiles and dynamic contexts without manual reconfiguration?
  • RQ4Does the use of LLMs improve the adaptability and responsiveness of digital health interventions in real-time behavioral support?

Key findings

  • GPT-4-generated JITAIs significantly outperformed both laypersons and healthcare professionals in overall quality across all evaluation metrics: appropriateness, engagement, effectiveness, and professionalism.
  • The LLM demonstrated consistent performance across all three patient personas, maintaining high relevance and personalization regardless of clinical severity or context.
  • In terms of engagement, GPT-4 messages were rated as more motivating and actionable than human-generated alternatives by both expert and lay raters.
  • The system achieved high contextual accuracy, with messages dynamically adapting to real-time factors such as mood, activity level, and social environment.
  • The results suggest that LLMs can reduce the burden of manual JITAI authoring while improving scalability and personalization in digital health interventions.
  • GPT-4’s output was rated as clinically sound and aligned with evidence-based cardiac rehabilitation guidelines, indicating strong potential for integration into clinical workflows.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.