[Paper Review] Understanding the Therapeutic Relationship between Counselors and Clients in Online Text-based Counseling using LLMs
This paper proposes an LLM-based automatic evaluation framework for assessing the working alliance in online text-based counseling, using comprehensive guidelines and Chain-of-Thought prompting to align LLM outputs with human expert judgments. The approach achieves high agreement with human evaluations and demonstrates potential for LLMs to serve as supervisory tools in mental health counseling.
Robust therapeutic relationships between counselors and clients are fundamental to counseling effectiveness. The assessment of therapeutic alliance is well-established in traditional face-to-face therapy but may not directly translate to text-based settings. With millions of individuals seeking support through online text-based counseling, understanding the relationship in such contexts is crucial. In this paper, we present an automatic approach using large language models (LLMs) to understand the development of therapeutic alliance in text-based counseling. We adapt a theoretically grounded framework specifically to the context of online text-based counseling and develop comprehensive guidelines for characterizing the alliance. We collect a comprehensive counseling dataset and conduct multiple expert evaluations on a subset based on this framework. Our LLM-based approach, combined with guidelines and simultaneous extraction of supportive evidence underlying its predictions, demonstrates effectiveness in identifying the therapeutic alliance. Through further LLM-based evaluations on additional conversations, our findings underscore the challenges counselors face in cultivating strong online relationships with clients. Furthermore, we demonstrate the potential of LLM-based feedback mechanisms to enhance counselors' ability to build relationships, supported by a small-scale proof-of-concept.
Motivation & Objective
- To address the high cost and subjectivity of manual evaluation in online mental health counseling.
- To develop an automatic, third-party evaluation method for the working alliance that is impartial and scalable.
- To improve LLM performance in assessing counseling quality by using expert-designed guidelines and Chain-of-Thought prompting.
- To validate the LLM-based evaluation against human-annotated data and assess its reliability and interpretability.
- To explore the potential of LLMs as supervisory tools for psychotherapists in clinical training and supervision.
Proposed method
- The authors collected a large-scale text-based counseling dataset from an online platform, including self-reported working alliance scores from both counselors and clients.
- They developed an observer version of the Working Alliance Inventory (WAI) based on Bordin’s therapeutic relationship theory, with four detailed questions per component (goal, task, bond).
- Expert annotators labeled a subset of sessions with evidence-based assessments across five levels: considerable evidence against, some against, no evidence against, some for, and considerable evidence for.
- LLMs such as GPT-4 were fine-tuned using these guidelines to improve their accuracy in evaluating the working alliance across entire multi-turn conversations.
- Chain-of-Thought (CoT) prompting was integrated to enhance LLMs’ ability to extract and justify evidence from the conversation, improving interpretability and consistency.
- The framework was validated through comparison with human annotations, showing high agreement and improved annotation consistency when using LLM-extracted evidence.
Experimental results
Research questions
- RQ1Can LLMs be effectively guided by expert-designed guidelines to evaluate the working alliance in online counseling with high reliability?
- RQ2How does Chain-of-Thought prompting improve the interpretability and accuracy of LLM-based counseling evaluations?
- RQ3To what extent does LLM-based evaluation align with human expert assessments of therapeutic alliance?
- RQ4Can LLM-extracted evidence improve the consistency and quality of human annotation in counseling evaluation?
- RQ5What is the potential of LLMs as supervisory tools in mental health counseling training and quality assurance?
Key findings
- The LLM-based evaluation method achieved high agreement with human expert evaluations, demonstrating strong reliability and validity in assessing the working alliance.
- The integration of expert-designed guidelines significantly improved GPT-4’s performance in evaluating counseling sessions, ensuring internal consistency and alignment with human judgments.
- Chain-of-Thought prompting enabled the LLM to identify and extract supportive evidence from conversations, enhancing interpretability and justifiability of scores.
- LLM-extracted evidence improved the consistency among human annotators, suggesting that LLMs can serve as valuable tools in human evaluation processes.
- The approach provides a cost-effective, scalable, and impartial alternative to traditional manual evaluation, with strong potential for use in clinical supervision and training.
- The study confirms that LLMs can be effectively leveraged for automatic, third-party assessment of counseling quality, particularly in relational dynamics like the working alliance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.