[Paper Review] LLM-based Conversational AI Therapist for Daily Functioning Screening and Psychotherapeutic Intervention via Everyday Smart Devices
CaiTI is an LLM-driven conversational AI therapist that screens 37 day-to-day functioning dimensions, provides empathic validation, and delivers MI and CBT-based interventions via common smart devices, with real-world pilot deployment.
Despite the global mental health crisis, access to screenings, professionals, and treatments remains high. In collaboration with licensed psychotherapists, we propose a Conversational AI Therapist with psychotherapeutic Interventions (CaiTI), a platform that leverages large language models (LLM)s and smart devices to enable better mental health self-care. CaiTI can screen the day-to-day functioning using natural and psychotherapeutic conversations. CaiTI leverages reinforcement learning to provide personalized conversation flow. CaiTI can accurately understand and interpret user responses. When the user needs further attention during the conversation, CaiTI can provide conversational psychotherapeutic interventions, including cognitive behavioral therapy (CBT) and motivational interviewing (MI). Leveraging the datasets prepared by the licensed psychotherapists, we experiment and microbenchmark various LLMs' performance in tasks along CaiTI's conversation flow and discuss their strengths and weaknesses. With the psychotherapists, we implement CaiTI and conduct 14-day and 24-week studies. The study results, validated by therapists, demonstrate that CaiTI can converse with users naturally, accurately understand and interpret user responses, and provide psychotherapeutic interventions appropriately and effectively. We showcase the potential of CaiTI LLMs to assist the mental therapy diagnosis and treatment and improve day-to-day functioning screening and precautionary psychotherapeutic intervention systems.
Motivation & Objective
- Screen day-to-day functioning across 37 dimensions to assess mental health status.
- Provide empathic validations and psychotherapeutic interventions tailored to physical and mental status.
- Leverage reinforcement learning to personalize the conversational flow and prioritize dimensions.
- Incorporate MI and CBT within a natural, therapist-guided conversation.
- Validate performance through therapist-informed microbenchmarks and real-world deployments.
Proposed method
- Utilize LLMs to generate open-ended questions and semantically analyze user responses within a 37-dimension framework.
- Apply Q-learning to guide the next question with a 39-state space and an epsilon-greedy policy.
- Use a Response Analyzer to segment user input and classify segments into (Dimension, Score) across 37 dimensions and 3 scores.
- Incorporate a reflection-validation (R-V) process with MI techniques to handle high-salience dimensions.
- End each session with a four-step CBT process to address identified issues.
- Develop task-specific Reasoners, Guides, and Validators to ensure quality and reduce AI biases during psychotherapy.
Experimental results
Research questions
- RQ1Can CaiTI accurately screen day-to-day functioning across 37 dimensions using natural, open-ended conversations?
- RQ2Do LLM-based modules can effectively generate, analyze, and route psychotherapeutic interventions (MI and CBT) in real-time?
- RQ3How does reinforcement learning influence the personalization and prioritization of screening questions?
- RQ4What is the comparative performance of different LLMs in the response analysis and psychotherapy components?
- RQ5Is CaiTI feasible and acceptable in real-world, long-term deployments with therapeutic validation by licensed psychotherapists?
Key findings
- CaiTI demonstrates natural conversational ability and accurate interpretation of user responses for delivering psychotherapeutic interventions.
- Therapist-validated studies show CaiTI can provide appropriate and effective psychotherapeutic interventions.
- Task-specific LLMs (Reasoner, Guide, Validator) help mitigate biases and improve psychotherapy quality.
- Microbenchmarks indicate GPT-4 and GPT-3.5 Turbo offer strong performance for the Response Analyzer task, with Llama-2 models underperforming on dimension/score classification.
- A 14-day to 24-week real-world deployment with 20 subjects provides evidence of CaiTI’s capability to assess status and deliver interventions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.