[논문 리뷰] Large Language Models for Mental Health Diagnostic Assessments: Exploring The Potential of Large Language Models for Assisting with Mental Health Diagnostic Assessments -- The Depression and Anxiety Case
논문은 PHQ-9 및 GAD-7 진단 평가를 돕기 위해 LLM의 프롬프트 및 파인튜닝을 평가하고, 독점 모델과 오픈소스 모델을 전문가 ground truth와 비교하며, DiagnosticLlama 모델 및 관련 데이터셋을 공개합니다.
Large language models (LLMs) are increasingly attracting the attention of healthcare professionals for their potential to assist in diagnostic assessments, which could alleviate the strain on the healthcare system caused by a high patient load and a shortage of providers. For LLMs to be effective in supporting diagnostic assessments, it is essential that they closely replicate the standard diagnostic procedures used by clinicians. In this paper, we specifically examine the diagnostic assessment processes described in the Patient Health Questionnaire-9 (PHQ-9) for major depressive disorder (MDD) and the Generalized Anxiety Disorder-7 (GAD-7) questionnaire for generalized anxiety disorder (GAD). We investigate various prompting and fine-tuning techniques to guide both proprietary and open-source LLMs in adhering to these processes, and we evaluate the agreement between LLM-generated diagnostic outcomes and expert-validated ground truth. For fine-tuning, we utilize the Mentalllama and Llama models, while for prompting, we experiment with proprietary models like GPT-3.5 and GPT-4o, as well as open-source models such as llama-3.1-8b and mixtral-8x7b.
연구 동기 및 목표
- LLMs가 표준화된 PHQ-9 및 GAD-7 진단 절차를 따를 수 있는지 평가.
- 프롬프트 방식과 파인튜닝 방식의 효과를 독점 모델과 오픈소스 모델 전반에서 비교.
- 진단 기준에 대해 파인튜닝된 특화 DiagnosticLlama 모델을 개발하고 평가.
- 연구 지원을 위해 임상의 주석이 달린 합성 데이터와 모델 산출물을 생성 및 공개.
제안 방법
- 모델 지침으로 PRIMATE의 PHQ-9 및 GAD-7 실제 데이터셋 사용.
- hits@k 및 표준 분류 지표(정확도, 정밀도, 재현율, F1)로 LLM 출력 평가.
- 프롬프트 탐색(무작위, 사례 기반, 가이드 기반)과 파인튜닝(SFT, RLHF, DPO)을 모델 전반에 걸쳐 탐색.
- DiagnosticLlama를 만들기 위해 MentalllaMa를 파인튜닝하고 프롬프트 결과와 비교.
- DiagnosticLlama 및 주석된 데이터셋을 Hugging Face와 GitHub를 통해 공개.

실험 결과
연구 질문
- RQ1LLMs가 게시물에서 PHQ-9 및 GAD-7 증상 기준을 식별하여 전문가 ground truth와 일치시킬 수 있는가?
- RQ2프롬프트와 파인튜닝이 임상의 평가와의 정렬에 어떤 영향을 미치는가?
- RQ3독점 모델과 오픈소스 모델이 진단 기준 준수에서 어떻게 비교되는가?
- RQ4신뢰할 수한 LLM 보조 정신건강 진단의 실용적 한계와 데이터 요구사항은 무엇인가?
주요 결과
- LLMs가 프롬프트와 파인튜닝 설정 모두에서 PHQ-9 및 GAD-7 작업에 대한 전문가 주석 품질에 접근한다.
- GPT-4o-mini와 mixtral-8x7b는 각각 독점 및 오픈소스 모델 중 핵심 평가에서 뛰어나다.
- DiagnosticLlama 모델의 파인튜닝은 유망한 결과를 보이나 이 작업은 자원 집약적이고 도전적이다.
- 구형 LLMs 및 비자회귀형 모델은 현대 LLM 대비 현저한 성능 격차를 보인다.
- Few-shot 프롬프트와 파인튜닝은 일반적으로 제로샷 대비 성능을 향상시킨다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.