Skip to main content
QUICK REVIEW

[논문 리뷰] Comparing the Efficacy of GPT-4 and Chat-GPT in Mental Health Care: A Blind Assessment of Large Language Models for Psychological Support

Birger Moëll|arXiv (Cornell University)|2024. 05. 15.
Artificial Intelligence in Healthcare and Education인용 수 4
한 줄 요약

이 연구는 정신건강 전문의가 이끄는 盲검사 방식으로, GPT-4와 ChatGPT가 우울증, 불안, 외상 및 관련 장애를 포함한 18개의 심리적 프롬프트에 대해 어떻게 반응하는지 비교 평가한다. GPT-4는 ChatGPT보다 유의미하게 뛰어난 성과를 보였으며, 평균 임상 평가 점수는 8.29/10에 달했고, ChatGPT는 6.52/10이었다. 이는 공감 능력, 임상적 관련성 및 치료 지침 면에서 뛰어난 성능을 보임을 시사한다.

ABSTRACT

Background: Rapid advancements in natural language processing have led to the development of large language models with the potential to revolutionize mental health care. These models have shown promise in assisting clinicians and providing support to individuals experiencing various psychological challenges. Objective: This study aims to compare the performance of two large language models, GPT-4 and Chat-GPT, in responding to a set of 18 psychological prompts, to assess their potential applicability in mental health care settings. Methods: A blind methodology was employed, with a clinical psychologist evaluating the models' responses without knowledge of their origins. The prompts encompassed a diverse range of mental health topics, including depression, anxiety, and trauma, to ensure a comprehensive assessment. Results: The results demonstrated a significant difference in performance between the two models (p > 0.05). GPT-4 achieved an average rating of 8.29 out of 10, while Chat-GPT received an average rating of 6.52. The clinical psychologist's evaluation suggested that GPT-4 was more effective at generating clinically relevant and empathetic responses, thereby providing better support and guidance to potential users. Conclusions: This study contributes to the growing body of literature on the applicability of large language models in mental health care settings. The findings underscore the importance of continued research and development in the field to optimize these models for clinical use. Further investigation is necessary to understand the specific factors underlying the performance differences between the two models and to explore their generalizability across various populations and mental health conditions.

연구 동기 및 목표

  • 면허를 취득한 정신건강 전문의가 이끄는 막힌 평가를 통해 GPT-4와 ChatGPT가 심리적 지원을 제공하는 데 있어 임상적 효과를 평가하는 것.
  • 우울증, 불안, 외상, 정서 조절 등 다양한 정신건강 주제에서 두 대규모 언어 모델의 반응 품질을 비교하는 것.
  • LLM이 정신건강 치료 환경에 적합한 임상적으로 관련성 있고 공감 능력 있는, 목표 지향적인 반응을 제공할 수 있는지 평가하는 것.
  • 성능 차이와 안전성 고려사항을 규명하여 LLM의 윤리적 통합을 정신건강 지원에 도와주는 것.

제안 방법

  • 면밀한 평가 설계를 통해, 임상 심리학자가 모델의 출처를 알지 못한 채 반응을 평가하였다.
  • 18개의 표준화된 심리적 프롬프트를 사용하였으며, 불안, 우울증, 자존감, 스트레스, 인간관계 문제 등 핵심 영역을 포함하였다.
  • 임상적 관련성, 공감 능력, 치료 정확도, 목표 지향성 기준으로 10점 만점 척도로 평가하였다.
  • GPT-4와 ChatGPT 간 평균 반응 점수를 비교하기 위해 통계 분석을 실시하였으며, 유의미성은 p < 0.05로 설정하였다.
  • 언어 톤, 정서적 반응성, 임상 적합성에 대한 정성적 및 정량적 평가에 중점을 두었다.
Figure 1 : Average rating of psychological advice generated by GPT-4 / Chat-GPT.
Figure 1 : Average rating of psychological advice generated by GPT-4 / Chat-GPT.

실험 결과

연구 질문

  • RQ1GPT-4와 ChatGPT는 심리적 프롬프트에 대해 공감 능력 있고 임상적으로 관련성 있는 반응을 어떻게 제공하는가?
  • RQ2어느 모델이 우울증, 불안 및 외상 증상에 대해 더 강한 치료적 일치를 보이는가?
  • RQ3LLM은 임상적 감독 없이 목표 지향적인 심리적 간호를 어느 정도 지원할 수 있는가?
  • RQ4두 모델 간 반응 구조, 정서적 톤, 임상적 유용성에서의 주요 정성적 차이점은 무엇인가?

주요 결과

  • GPT-4는 평균 임상 평가 점수 8.29/10을 기록하여 ChatGPT의 6.52/10보다 유의미하게 높았다 (p < 0.05).
  • GPT-4는 더 뛰어난 공감 능력을 보였으며, 사용자 고통에 더 온전히 공감하고 정서적으로 연결된 반응을 제공했다.
  • GPT-4는 우울증, 불안, 강박장애(OCD) 등의 증상 관리에 더 정확하고 실행 가능한 전략을 제공하였다.
  • ChatGPT의 반응는 더 덜 뉘앙스가 있고, 일부는 명확한 치료 방향이 없어 안전성 측면에서 떨어졌다.
  • 임상 심리학자는 GPT-4가 언어와 어조를 통해 더 효과적으로 치료적 유대감을 형성한다고 평가했다.
  • 성능 차이는 외상, ADHD, 사회성 불안과 관련된 다양한 프롬프트 전반에서 일관되게 나타났다.
Figure 2 : Rating for each response generated by GPT-4 / Chat-GPT.
Figure 2 : Rating for each response generated by GPT-4 / Chat-GPT.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.