[Paper Review] Toward Large Language Models as a Therapeutic Tool: Comparing Prompting Techniques to Improve GPT-Delivered Problem-Solving Therapy
This study evaluates prompt engineering techniques to enhance GPT-4’s ability to deliver Problem-Solving Therapy (PST) for family caregivers, using in-context learning with few-shot examples and Chain-of-Thought prompting. It finds that few-shot prompting significantly improves symptom identification and goal setting, while Chain-of-Thought boosts empathy but reduces task accuracy, demonstrating that prompt design critically shapes LLM performance in therapeutic dialogue.
While Large Language Models (LLMs) are being quickly adapted to many domains, including healthcare, their strengths and pitfalls remain under-explored. In our study, we examine the effects of prompt engineering to guide Large Language Models (LLMs) in delivering parts of a Problem-Solving Therapy (PST) session via text, particularly during the symptom identification and assessment phase for personalized goal setting. We present evaluation results of the models' performances by automatic metrics and experienced medical professionals. We demonstrate that the models' capability to deliver protocolized therapy can be improved with the proper use of prompt engineering methods, albeit with limitations. To our knowledge, this study is among the first to assess the effects of various prompting techniques in enhancing a generalist model's ability to deliver psychotherapy, focusing on overall quality, consistency, and empathy. Exploring LLMs' potential in delivering psychotherapy holds promise with the current shortage of mental health professionals amid significant needs, enhancing the potential utility of AI-based and AI-enhanced care services.
Motivation & Objective
- To assess whether prompt engineering can enhance a general-purpose LLM’s ability to deliver protocolized Problem-Solving Therapy (PST) without fine-tuning.
- To evaluate the impact of different prompting techniques—zero-shot, few-shot, and Chain-of-Thought—on the quality, consistency, and empathy of LLM-generated PST dialogues.
- To compare LLM-generated dialogues against a human therapist baseline using expert evaluations and automated metrics.
- To explore the feasibility of using off-the-shelf LLMs as therapeutic tools in mental health care, particularly for addressing provider shortages.
- To identify which prompting strategies best balance therapeutic quality, task accuracy, and empathetic responsiveness in LLM-driven psychotherapy.
Proposed method
- Adapted an existing LLM pipeline for medical Q&A to generate dialogues for PST, focusing on symptom identification and goal setting.
- Employed in-context learning via few-shot prompting, providing the model with example dialogues to guide response style and structure.
- Applied Chain-of-Thought (CoT) prompting to encourage step-by-step reasoning, aiming to improve empathy and exploration in responses.
- Used standardized patient actors to simulate consistent, persona-based interactions, ensuring reproducible evaluation conditions.
- Conducted dual evaluation: automated metrics for coherence and structure, and expert human ratings for empathy, quality, and consistency.
- Compared model outputs against a rule-based chatbot baseline, which used generic empathetic phrases instead of evidence-based therapeutic techniques.
Experimental results
Research questions
- RQ1How do different prompting techniques (zero-shot, few-shot, Chain-of-Thought) affect the quality and consistency of LLM-generated Problem-Solving Therapy dialogues?
- RQ2To what extent can few-shot prompting improve the LLM’s accuracy in symptom identification and goal setting during PST sessions?
- RQ3Does Chain-of-Thought prompting enhance empathy in LLM responses, and if so, at what cost to task-specific performance?
- RQ4How does the LLM’s performance compare to a human therapist baseline in a controlled, expert-evaluated setting?
- RQ5What are the limitations of prompt engineering in enabling LLMs to deliver protocolized, patient-centered psychotherapy without model fine-tuning?
Key findings
- Few-shot prompting significantly improved the LLM’s performance in symptom identification and goal setting compared to zero-shot prompting.
- Chain-of-Thought prompting increased perceived empathy, particularly in the exploration phase of therapy, but reduced accuracy in symptom identification and goal setting.
- The LLM-generated dialogues received higher expert ratings for overall quality than the human-created baseline, which used generic empathetic phrases instead of evidence-based techniques.
- Despite improvements, the model still struggled with implicit therapeutic principles, such as favoring actionable advice over overly optimistic or generic responses.
- The results suggest that prompt engineering can meaningfully enhance LLMs for therapeutic applications, but the trade-offs between empathy and task accuracy require careful design.
- The study highlights the risk of LLMs regurgitating private information if not properly safeguarded, underscoring the need for privacy-preserving fine-tuning and retrieval-augmented generation in future work.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.