[Paper Review] Large Language Models in Mental Health Care: a Scoping Review
This scoping review analyzes 34 studies on large language models in mental health care to map applications, datasets, training methods, ethics, and validation gaps.
Objectieve:This review aims to deliver a comprehensive analysis of Large Language Models (LLMs) utilization in mental health care, evaluating their effectiveness, identifying challenges, and exploring their potential for future application. Materials and Methods: A systematic search was performed across multiple databases including PubMed, Web of Science, Google Scholar, arXiv, medRxiv, and PsyArXiv in November 2023. The review includes all types of original research, regardless of peer-review status, published or disseminated between October 1, 2019, and December 2, 2023. Studies were included without language restrictions if they employed LLMs developed after T5 and directly investigated research questions within mental health care settings. Results: Out of an initial 313 articles, 34 were selected based on their relevance to LLMs applications in mental health care and the rigor of their reported outcomes. The review identified various LLMs applications in mental health care, including diagnostics, therapy, and enhancing patient engagement. Key challenges highlighted were related to data availability and reliability, the nuanced handling of mental states, and effective evaluation methods. While LLMs showed promise in improving accuracy and accessibility, significant gaps in clinical applicability and ethical considerations were noted. Conclusion: LLMs hold substantial promise for enhancing mental health care. For their full potential to be realized, emphasis must be placed on developing robust datasets, development and evaluation frameworks, ethical guidelines, and interdisciplinary collaborations to address current limitations.
Motivation & Objective
- Survey dataset types, models, training techniques, and their suitability for mental health tasks.
- Characterize mental health applications enabled by LLMs (diagnosis, therapy, engagement, screening, education).
- Identify validation measures, performance metrics, and evaluation practices.
- Examine ethical, privacy, safety, and regulatory challenges in deploying LLMs for mental health care.
- Highlight gaps between current tools and clinical practicality to guide future work.
Proposed method
- Adheres to PRISMA 2020 guidelines for scoping reviews.
- Comprehensive search across PubMed, Web of Science, Google Scholar, arXiv, medRxiv, PsyArXiv conducted in Nov 2023.
- Initial identification of 313 publications; 34 met inclusion criteria after screening.
- GPT-4 assisted in title/abstract screening as a secondary reviewer with Cohen’s Kappa ≈ 0.90 against a human reviewer.
- Publications categorized into Dataset/Benchmark, Model Development/Fine-tuning, Application/Evaluation, and Ethics/Safety.
- Distinction between prompting-based and fine-tuned LLMs; emphasis on instruction fine-tuning (IFT) and prompt-tuning strategies.

Experimental results
Research questions
- RQ1What datasets and models are used for mental health tasks with LLMs?
- RQ2What mental health applications are addressed by LLMs and how are they validated?
- RQ3What are the ethical, privacy, safety, and governance considerations for LLMs in mental health care?
- RQ4What gaps exist between current LLM tools and clinical practicality, and what is needed to bridge them?
Key findings
- LLMs are applied to conversational agents, empathetic dialogue, screening, and supportive tools for both patients and clinicians.
- Most studies use 2022–2023 publications, with a surge in prompt-tuning and application-focused work; few dataset/benchmark papers.
- Datasets are largely drawn from social media, with some clinician-generated dialogues and synthetic data; licensing is often non-commercial.
- Evaluation relies heavily on automation metrics like F1, accuracy, recall, and precision, with limited standardized clinical validation.
- Ethical, privacy, and safety concerns are underexplored, signaling a need for robust data governance and interdisciplinary collaboration.
- Overall, LLMs show potential for diagnostics and patient support, but clinical practicality and ethical integration require further development.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.