[Paper Review] Comparing the Efficacy of GPT-4 and Chat-GPT in Mental Health Care: A Blind Assessment of Large Language Models for Psychological Support
This study conducts a blind, clinician-led evaluation comparing GPT-4 and ChatGPT in responding to 18 psychological prompts across depression, anxiety, trauma, and related conditions. GPT-4 significantly outperformed ChatGPT, achieving an average clinical rating of 8.29/10 versus 6.52/10, demonstrating superior empathy, clinical relevance, and therapeutic guidance.
Background: Rapid advancements in natural language processing have led to the development of large language models with the potential to revolutionize mental health care. These models have shown promise in assisting clinicians and providing support to individuals experiencing various psychological challenges. Objective: This study aims to compare the performance of two large language models, GPT-4 and Chat-GPT, in responding to a set of 18 psychological prompts, to assess their potential applicability in mental health care settings. Methods: A blind methodology was employed, with a clinical psychologist evaluating the models' responses without knowledge of their origins. The prompts encompassed a diverse range of mental health topics, including depression, anxiety, and trauma, to ensure a comprehensive assessment. Results: The results demonstrated a significant difference in performance between the two models (p > 0.05). GPT-4 achieved an average rating of 8.29 out of 10, while Chat-GPT received an average rating of 6.52. The clinical psychologist's evaluation suggested that GPT-4 was more effective at generating clinically relevant and empathetic responses, thereby providing better support and guidance to potential users. Conclusions: This study contributes to the growing body of literature on the applicability of large language models in mental health care settings. The findings underscore the importance of continued research and development in the field to optimize these models for clinical use. Further investigation is necessary to understand the specific factors underlying the performance differences between the two models and to explore their generalizability across various populations and mental health conditions.
Motivation & Objective
- To evaluate the clinical efficacy of GPT-4 and ChatGPT in providing psychological support through a blind assessment by a licensed psychologist.
- To compare the quality of responses from two large language models across diverse mental health topics, including depression, anxiety, trauma, and emotional regulation.
- To assess whether LLMs can deliver clinically relevant, empathetic, and goal-directed responses suitable for mental health care settings.
- To inform ethical integration of LLMs in mental health support by identifying performance differences and safety considerations.
Proposed method
- A blind evaluation design was employed, where a clinical psychologist rated model responses without knowledge of the model origin.
- Eighteen standardized psychological prompts were used, covering key areas such as anxiety, depression, self-esteem, stress, and relationship issues.
- Responses were scored on a 10-point scale based on clinical relevance, empathy, therapeutic accuracy, and goal-directedness.
- Statistical analysis was conducted to compare mean response scores between GPT-4 and ChatGPT, with p < 0.05 indicating significance.
- The evaluation focused on qualitative and quantitative assessment of linguistic tone, emotional responsiveness, and clinical appropriateness.

Experimental results
Research questions
- RQ1How do GPT-4 and ChatGPT compare in delivering empathetic and clinically relevant responses to psychological prompts?
- RQ2Which model demonstrates stronger therapeutic alignment in addressing symptoms of depression, anxiety, and trauma?
- RQ3To what extent do LLMs support goal-directed psychological interventions without clinical oversight?
- RQ4What are the key qualitative differences in response structure, emotional tone, and clinical utility between the two models?
Key findings
- GPT-4 achieved a mean clinical rating of 8.29 out of 10, significantly higher than ChatGPT’s 6.52 (p < 0.05).
- GPT-4 demonstrated superior empathy, with responses perceived as more compassionate and emotionally attuned to user distress.
- GPT-4 provided more clinically accurate and actionable strategies for managing conditions like depression, anxiety, and OCD.
- ChatGPT responses were rated as less nuanced and occasionally less safe, with some lacking clear therapeutic direction.
- The clinical psychologist judged GPT-4 as more effective in fostering a therapeutic alliance through language and tone.
- Performance differences were consistent across diverse prompts, including those related to trauma, ADHD, and social anxiety.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.