[Paper Review] GPTCoach: Towards LLM-Based Physical Activity Coaching
GPTCoach is an LLM-based chatbot that implements an evidence-based health coaching program, uses motivational interviewing strategies, and can query wearable health data to support physical activity behavior change; evaluated as a technology probe with 16 participants.
Mobile health applications show promise for scalable physical activity promotion but are often insufficiently personalized. In contrast, health coaching offers highly personalized support but can be prohibitively expensive and inaccessible. This study draws inspiration from health coaching to explore how large language models (LLMs) might address personalization challenges in mobile health. We conduct formative interviews with 12 health professionals and 10 potential coaching recipients to develop design principles for an LLM-based health coach. We then built GPTCoach, a chatbot that implements the onboarding conversation from an evidence-based coaching program, uses conversational strategies from motivational interviewing, and incorporates wearable data to create personalized physical activity plans. In a lab study with 16 participants using three months of historical data, we find promising evidence that GPTCoach gathers rich qualitative information to offer personalized support, with users feeling comfortable sharing concerns. We conclude with implications for future research on LLM-based physical activity support.
Motivation & Objective
- Identify how health experts coach to overcome barriers to physical activity and what LLMs could contribute to these strategies.
- Assess how self-tracking data is used to promote activity and how LLMs might leverage such data for coaching.
- Design a facilitative, non-prescriptive AI coach grounded in an established coaching program and MI techniques.
- Evaluate GPTCoach’s adherence to coaching principles and its use of data in real conversations.
Proposed method
- Conduct formative interviews with 12 health experts and 10 non-experts to extract design considerations for an LLM health coach.
- Develop GPTCoach as an onboarding conversation aligned with a validated health coaching program and motivational interviewing techniques.
- Implement a data and prompting pipeline with tool calls to fetch wearable data via HealthKit and visualize it in the UI.
- Employ prompt chaining to ensure adherence to the coaching program, MI strategies, and appropriate data use.
- Prototype and pilot-test with 16 participants to assess MI behavior, coaching adherence, and data utilization.
Experimental results
Research questions
- RQ1RQ1: What coaching strategies do health experts use and which could LLMs adopt to overcome barriers to physical activity?
- RQ2RQ2: How do health experts use self-tracking data, and how might LLMs leverage this data to promote activity?
- RQ3RQ3: Can an LLM-based coach maintain a facilitative, non-judgmental coaching style while integrating personal data?
- RQ4RQ4: How effectively can prompt chaining enforce adherence to a structured coaching program in an LLM?
- RQ5RQ5: What are the risks and limitations of using LLMs for health coaching in terms of data usage and personalization?
Key findings
- MI-consistent or neutral behaviors occurred 84% of the time in automated MI coding.
- Participants reported feeling supported and comfortable sharing concerns with the chatbot.
- Prompt chaining helped GPTCoach adhere to the coaching program and initiate appropriate tool calls.
- Data use by GPTCoach was more variable, with some conversations leveraging data for motivation and others not proactively integrating data.
- Compared to vanilla GPT-4, GPTCoach showed greater alignment with MI principles, asking more open questions and giving less unsolicited advice.
- Participants believed AI could augment data analysis for goal setting and accountability, but privacy and personalization challenges remained.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.