[Paper Review] Supplementary Material for "The Unreasonable Effectiveness of Large Language Models in Issuing Just-in-Time Adaptive Interventions: Fostering Physical Activity in a Prospective Cardiac Rehabilitation Setting"
This study evaluates the use of GPT-4 to generate personalized, context-aware Just-in-Time Adaptive Interventions (JITAIs) for promoting physical activity in cardiac rehabilitation. Using 450 patient-specific interventions across diverse personas and contexts, GPT-4 outperformed both laypersons and healthcare professionals in appropriateness, engagement, effectiveness, and professionalism, demonstrating LLMs' potential to revolutionize scalable, adaptive digital health support.
This upload contains supplementary materials for the publication: "The Last JITAI? - The Unreasonable Effectiveness of Large Language Models in Issuing Just-in-Time Adaptive Interventions: Fostering Physical Activity in a Prospective Cardiac Rehabilitation Setting" Supplement 1: Persona and context information used as basis for JITAI decision-making and content generation Supplement 2: Instructions given to JITAI generators Supplement 3: Data and analysis skripts (contains a Read_me file specifying further details)
Motivation & Objective
- To assess the feasibility and quality of Large Language Models (LLMs) in generating Just-in-Time Adaptive Interventions (JITAIs) for physical activity promotion in cardiac rehabilitation.
- To address scalability and flexibility limitations of traditional JITAI implementation models that rely on manual or rule-based design.
- To compare LLM-generated JITAI suggestions against those created by laypersons and healthcare professionals in real-world clinical relevance and quality.
- To evaluate the performance of GPT-4 in generating context-sensitive, patient-tailored behavioral health messages across varying clinical severities and situational contexts.
Proposed method
- The study used GPT-4 to generate 450 JITAI decisions and messages based on five distinct context sets per patient persona, representing varying levels of cardiovascular disease severity.
- Three patient personas were defined, each reflecting different clinical profiles and behavioral needs in cardiac rehabilitation.
- Context sets included real-time factors such as activity level, mood, social context, and clinical status, simulating dynamic, real-world conditions.
- Generated messages were systematically evaluated against human-generated interventions from 10 laypersons and 10 healthcare professionals using four key metrics: appropriateness, engagement, effectiveness, and professionalism.
- Evaluation was conducted via expert and user-based assessment, ensuring alignment with clinical guidelines and patient-centered design principles.
- The LLM was prompted with structured templates and clinical context to ensure consistency and relevance in message generation.
Experimental results
Research questions
- RQ1Can GPT-4 generate JITAI messages that are as clinically appropriate and engaging as those created by healthcare professionals in a cardiac rehabilitation context?
- RQ2How does the quality of LLM-generated JITAIs compare to those produced by laypersons in terms of personalization and clinical relevance?
- RQ3To what extent can LLMs scale personalized, context-aware interventions across diverse patient profiles and dynamic contexts without manual reconfiguration?
- RQ4Does the use of LLMs improve the adaptability and responsiveness of digital health interventions in real-time behavioral support?
Key findings
- GPT-4-generated JITAIs significantly outperformed both laypersons and healthcare professionals in overall quality across all evaluation metrics: appropriateness, engagement, effectiveness, and professionalism.
- The LLM demonstrated consistent performance across all three patient personas, maintaining high relevance and personalization regardless of clinical severity or context.
- In terms of engagement, GPT-4 messages were rated as more motivating and actionable than human-generated alternatives by both expert and lay raters.
- The system achieved high contextual accuracy, with messages dynamically adapting to real-time factors such as mood, activity level, and social environment.
- The results suggest that LLMs can reduce the burden of manual JITAI authoring while improving scalability and personalization in digital health interventions.
- GPT-4’s output was rated as clinically sound and aligned with evidence-based cardiac rehabilitation guidelines, indicating strong potential for integration into clinical workflows.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.