[Paper Review] A Computational Framework for Behavioral Assessment of LLM Therapists
Bolt is a framework to systematically characterize LLM therapist behavior, comparing it to high- and low-quality human therapy, and exploring how prompts affect behavior. It uses simulated client-LLM conversations and a psychotherapy-technique taxonomy to identify behaviors.
The emergence of large language models (LLMs) like ChatGPT has increased interest in their use as therapists to address mental health challenges and the widespread lack of access to care. However, experts have emphasized the critical need for systematic evaluation of LLM-based mental health interventions to accurately assess their capabilities and limitations. Here, we propose BOLT, a proof-of-concept computational framework to systematically assess the conversational behavior of LLM therapists. We quantitatively measure LLM behavior across 13 psychotherapeutic approaches with in-context learning methods. Then, we compare the behavior of LLMs against high- and low-quality human therapy. Our analysis based on Motivational Interviewing therapy reveals that LLMs often resemble behaviors more commonly exhibited in low-quality therapy rather than high-quality therapy, such as offering a higher degree of problem-solving advice when clients share emotions. However, unlike low-quality therapy, LLMs reflect significantly more upon clients' needs and strengths. Our findings caution that LLM therapists still require further research for consistent, high-quality care.
Motivation & Objective
- Motivate the need for systematic behavioral assessment of LLMs used in mental health care.
- Develop a computational framework (Bolt) to quantify LLM therapist behavior across a range of techniques.
- Compare LLM therapist behavior to high- and low-quality human therapy.
- Explore how prompting and model choice influence behavioral alignment with high-quality therapy.
Proposed method
- Introduce Bolt, a system prompts-based framework to simulate therapy conversations between LLMs and simulated clients using public therapy datasets.
- Annotate utterances with 13 therapist and 6 client behaviors drawn from established psychotherapy techniques.
- Evaluate GPT-3 and GPT-4 family models, plus Llama2 variants, on multi-label and binary-label behavior classification tasks.
- Use in-context learning with psychotherapy definitions and examples to identify behaviors; compare with high- and low-quality human therapy baselines.
- Analyze frequency, temporal order of behaviors, and adaptability across models; assess the effect of explicit prompting variations on behavior.
Experimental results
Research questions
- RQ1Can Bolt reliably identify therapist and client behaviors from therapy conversations?
- RQ2How do LLM therapist behaviors compare to high- and low-quality human therapy sessions?
- RQ3Do prompting strategies and model choice steer LLMs toward higher-quality therapeutic behaviors?
- RQ4Are LLMs more prone to problem-solving/solutions or reflective/normalize behaviors compared with humans?
- RQ5To what extent can LLMs reflect client needs and strengths similarly to high-quality therapy?
Key findings
- Prompting with psychotherapy definitions and examples yields the best macro-F1 for therapist behavior (57.7% macro-F1).
- Client behavior classification with prompting (binary-label) achieves the best macro-F1 (76.7%).
- LLM therapists show higher problem-solving behavior similar to low-quality human therapy, but also reflect client emotions and experiences more than typical low-quality therapy.
- GPT-4 and GPT-3.5-turbo generally exhibit more solution-focused behavior than Llama2 variants, suggesting RLHF-aligned tendencies influence these patterns.
- Simulated LLM therapy often aligns more with low-quality human therapy in behavior frequencies, indicating current non-ideal alignment with high-quality care.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.