[Paper Review] LLM-empowered Chatbots for Psychiatrist and Patient Simulation: Application and Evaluation
This paper evaluates ChatGPT-powered doctor and patient chatbots for psychiatric diagnostic conversations, using iterative prompt design and human+automatic evaluation with psychiatrists and patients.
Empowering chatbots in the field of mental health is receiving increasing amount of attention, while there still lacks exploration in developing and evaluating chatbots in psychiatric outpatient scenarios. In this work, we focus on exploring the potential of ChatGPT in powering chatbots for psychiatrist and patient simulation. We collaborate with psychiatrists to identify objectives and iteratively develop the dialogue system to closely align with real-world scenarios. In the evaluation experiments, we recruit real psychiatrists and patients to engage in diagnostic conversations with the chatbots, collecting their ratings for assessment. Our findings demonstrate the feasibility of using ChatGPT-powered chatbots in psychiatric scenarios and explore the impact of prompt designs on chatbot behavior and user experience.
Motivation & Objective
- Formalize the task of doctor and patient chatbots for psychiatric outpatient diagnosis.
- Co-design prompts with psychiatrists to align chatbot behavior with real-world diagnostics.
- Develop and apply a human-centered evaluation framework combining user studies and automatic metrics.
- Demonstrate how prompt design affects chatbot empathy, in-depth questioning, and user experience.
Proposed method
- Iterative prompt design guided by collaborating psychiatrists to shape doctor and patient chatbots.
- Three-phase development: objective identification (Phase 1), prompt design and evaluation framework (Phase 2), and real-user evaluation with psychiatrists and patients (Phase 3).
- Empirical comparison of multiple prompt variants (D1–D4 for doctors; P1–P2 for patients).
- Combination of human evaluation metrics (fluency, empathy, expertise, engagement for doctors; ressemblance, rationality for patients) with automatic metrics (diagnosis accuracy, symptom recall, in-depth ratio, Distinct-1, etc.).
- Incorporation of empathy and in-depth questioning in doctor prompts; explicit honesty and colloquial/life-context language for patient prompts.

Experimental results
Research questions
- RQ1Can ChatGPT-powered doctor and patient chatbots approximate real psychiatric diagnostic conversations?
- RQ2How do different prompt designs affect chatbot empathy, in-depth questioning, and user experience?
- RQ3What is the relationship between human clinician-like behavior and automatic metrics in chatbot diagnostic tasks?
- RQ4How do simulated patients compare to real patients in terms of expressed symptoms and interaction style?
- RQ5What evaluation framework best captures the quality of psychiatric diagnostic dialogue?
Key findings
- Empathy-enabled doctor prompts improve empathy scores, but overly repetitive empathy can hurt user experience.
- A prompt design without explicit symptom-aspect prompts (D3) achieved the highest diagnostic accuracy among doctor chatbots in the study (55.56%), with variations in other metrics across prompts.
- Prompt variations influence both the depth of questioning and symptom recall, with D3 showing deeper in-depth questioning and higher symptom precision but lower symptom recall.
- Patient prompts that incorporate resistance and colloquial language (P2) yielded higher realism scores and better expression style from psychiatrists, though sometimes at the cost of unmentioned symptom ratio.
- Human doctors demonstrated more balanced and comprehensive symptom coverage than chatbots, highlighting remaining gaps in multi-disease screening and screening strategies.
- Automatic metrics revealed trade-offs between language style (Distinct-1) and accuracy in symptom reporting for patient chatbots.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.