Skip to main content
QUICK REVIEW

[Paper Review] Towards Explainable and Safe Conversational Agents for Mental Health: A Survey

Surjodeep Sarkar, Manas Gaur|arXiv (Cornell University)|Apr 25, 2023
Digital Mental Health InterventionsPsychology3 citations
TL;DR

This survey proposes a framework for developing explainable, safe, and trustworthy Virtual Mental Health Assistants (VMHAs) by integrating clinical knowledge, contextual understanding, and user-centered design. It emphasizes question/response generation grounded in clinical guidelines (e.g., PHQ-9), motivational interviewing, and diagnostic interviewing to improve triage accuracy and user trust, while addressing ethical and practical challenges in deployment.

ABSTRACT

Virtual Mental Health Assistants (VMHAs) are seeing continual advancements to support the overburdened global healthcare system that gets 60 million primary care visits, and 6 million Emergency Room (ER) visits annually. These systems are built by clinical psychologists, psychiatrists, and Artificial Intelligence (AI) researchers for Cognitive Behavioral Therapy (CBT). At present, the role of VMHAs is to provide emotional support through information, focusing less on developing a reflective conversation with the patient. A more comprehensive, safe and explainable approach is required to build responsible VMHAs to ask follow-up questions or provide a well-informed response. This survey offers a systematic critical review of the existing conversational agents in mental health, followed by new insights into the improvements of VMHAs with contextual knowledge, datasets, and their emerging role in clinical decision support. We also provide new directions toward enriching the user experience of VMHAs with explainability, safety, and wholesome trustworthiness. Finally, we provide evaluation metrics and practical considerations for VMHAs beyond the current literature to build trust between VMHAs and patients in active communications.

Motivation & Objective

  • To address the lack of contextual, explainable, and safe conversational agents in mental health by identifying critical gaps in current VMHA systems.
  • To propose a taxonomy of mental health conversations emphasizing user-level explainability and safety, particularly in active, reflective dialogue.
  • To integrate clinical knowledge (e.g., PHQ-9, DSM-V) and structured interviewing techniques (e.g., Motivational Interviewing, CDI) into VMHA design for improved decision support.
  • To improve VMHA trustworthiness through explainability, safety mechanisms, and ethical data handling, especially in high-stakes mental health triage.
  • To provide practical evaluation metrics and design considerations for VMHAs beyond current benchmarks, focusing on real-world usability and clinical integration.

Proposed method

  • Proposes a taxonomy of mental health conversations spanning low-level NLP (lexical, syntactic) to high-level discourse and pragmatics, with a focus on explainability and safety.
  • Integrates clinical guidelines (e.g., PHQ-9, DSM-V) into VMHA design to ground responses in evidence-based mental health assessment frameworks.
  • Advocates for knowledge-driven conversational agents that use structured question generation based on validated screening tools to detect mental health conditions.
  • Incorporates Motivational Interviewing (MI) techniques to enable empathetic, user-centered dialogue that resolves ambivalence and supports behavior change.
  • Proposes practical data curation strategies—such as anonymization, data abstraction, and synthetic conversation generation—to address privacy and annotation challenges.
  • Recommends evaluation metrics focused on clinical relevance, safety, and user trust, moving beyond standard NLP benchmarks to assess real-world impact.
Figure 1: Taxonomy of Mental Health Conversations: While connecting dots in our investigation from NLP-centered low-level analysis (lexical, morphological, syntactic, semantic) over Mental health conversations to the higher-level analysis (discourses, pragmatics), we determine the evaluation metrics
Figure 1: Taxonomy of Mental Health Conversations: While connecting dots in our investigation from NLP-centered low-level analysis (lexical, morphological, syntactic, semantic) over Mental health conversations to the higher-level analysis (discourses, pragmatics), we determine the evaluation metrics

Experimental results

Research questions

  • RQ1How can conversational agents in mental health be made more explainable and safe at the user level, especially in high-stakes triage scenarios?
  • RQ2To what extent do current VMHAs fail to integrate clinical knowledge (e.g., PHQ-9, DSM-V) in their response generation and question prompting?
  • RQ3How can Motivational Interviewing (MI) principles be adapted into VMHAs to improve user engagement and clinical outcomes?
  • RQ4What practical data and model design strategies can ensure privacy, quality, and trustworthiness in VMHA training and deployment?
  • RQ5What evaluation metrics and frameworks are needed to assess the clinical safety, explainability, and effectiveness of VMHAs beyond standard NLP benchmarks?

Key findings

  • Current VMHAs like Woebot, Wysa, and ChatGPT often respond to mental health queries without contextual grounding, leading to assumptive or irrelevant responses.
  • VMHAs that integrate clinical guidelines (e.g., PHQ-9) can detect mental health disturbances more accurately and trigger appropriate alerts to mental health professionals.
  • Knowledge-driven agents that use structured question generation based on validated tools show potential for improving triage reliability and clinical decision support.
  • The integration of Motivational Interviewing (MI) techniques enhances user engagement and supports behavior change by addressing ambivalence in a reflective, empathetic manner.
  • Ethical data practices—such as anonymization, data abstraction, and synthetic data generation—are critical for building trustworthy VMHAs while preserving user privacy.
  • Existing evaluation metrics in NLP are insufficient for assessing VMHA safety and explainability; new, clinically grounded metrics are needed for real-world deployment.
Figure 2: (Left) The outcome from existing VMHAs (e.g., WoeBot, Wysa) and ChatGPT (general purpose chatbot). (Right) Illustration of a knowledge-driven conversational agent in mental health (desired VMHA). The use of questions in PHQ-9 to induce conceptual flow in mental health conversational agents
Figure 2: (Left) The outcome from existing VMHAs (e.g., WoeBot, Wysa) and ChatGPT (general purpose chatbot). (Right) Illustration of a knowledge-driven conversational agent in mental health (desired VMHA). The use of questions in PHQ-9 to induce conceptual flow in mental health conversational agents

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.