Skip to main content
QUICK REVIEW

[Paper Review] Exploring The Design of Prompts For Applying GPT-3 based Chatbots: A Mental Wellbeing Case Study on Mechanical Turk

Harsh Kumar, Ilya Musabirov|arXiv (Cornell University)|Sep 22, 2022
AI in Service Interactions27 citations
TL;DR

The paper conducts a randomized factorial experiment (N=945) to study how prompt design (identity, intent, behavior) affects GPT-3 based chatbots for mood support, with both quantitative ratings and qualitative conversation analysis.

ABSTRACT

Large-Language Models like GPT-3 have the potential to enable HCI designers and researchers to create more human-like and helpful chatbots for specific applications. But evaluating the feasibility of these chatbots and designing prompts that optimize GPT-3 for a specific task is challenging. We present a case study in tackling these questions, applying GPT-3 to a brief 5-minute chatbot that anyone can talk to better manage their mood. We report a randomized factorial experiment with 945 participants on Mechanical Turk that tests three dimensions of prompt design to initialize the chatbot (identity, intent, and behaviour), and present both quantitative and qualitative analyses of conversations and user perceptions of the chatbot. We hope other HCI designers and researchers can build on this case study, for other applications of GPT-3 based chatbots to specific tasks, and build on and extend the methods we use for prompt design, and evaluation of the prompt design.

Motivation & Objective

  • Explore feasibility of GPT-3 chatbots for a brief mood-management task in a controlled setting.
  • Investigate how prompt modifiers (identity, intent, behavior) shape user perceptions and conversation dynamics.
  • Provide a methodological framework for prompt design exploration and evaluation in HCI.
  • Demonstrate data collection and analysis methods (quantitative surveys and qualitative logs) at a larger scale than pilot studies.

Proposed method

  • Use a factorial experimental design (2 x 3 x 3) to test 18 prompt arms combining identity, intent, and behavior modifiers.
  • Deploy a 5-minute open-ended mood-support chatbot to 945 MTurk participants with controlled interaction length and disclosure that it is AI.
  • Collect measures of perceived risk, trust, expertise, and willingness to interact again, plus conversation logs for qualitative analysis.
  • Conduct qualitative thematic analysis of participant comments and examine conversation dynamics across prompt conditions.
  • Analyze demographic and tech-affinity variables to understand heterogeneity in responses (e.g., prior mental health help, technology propensity).
  • Provide an open-source chat interface and methodology to enable replication and further prompt-engineering research.

Experimental results

Research questions

  • RQ1How do different prompt modifiers (identity, intent, behavior) influence user perceptions of risk, trust, expertise, and willingness to interact with a GPT-3 based chatbot?
  • RQ2What qualitative patterns emerge in 5-minute conversations when using GPT-3 chatbots for mood management across different prompt settings?
  • RQ3Do demographic factors or prior mental health help history modulate responses to GPT-3 chatbots under varying prompt designs?
  • RQ4Can factorial prompt design reveal main effects or interactions that guide better prompt engineering for task-specific chatbots?

Key findings

  • Participants showed moderately high perceived risk, moderate trust, high perceived expertise, and moderate willingness to interact again.
  • Past history of seeking mental health help significantly affected all four evaluation dimensions, with higher risk, lower trust, similar expertise, and higher willingness to interact again.
  • More favorable attitudes toward technology tended to associate with higher perceived risk, lower trust, higher expertise, and higher willingness to interact again.
  • Qualitative themes indicated generally positive interactions (about 70%), with data privacy concerns as a major negative theme (about 30%).
  • Friend versus Coach identity showed no clear qualitative/quantitative trend on conversation dynamics, but Friend tended to elicit slightly longer user responses; Intent appeared to influence word counts more than Identity or Behavior.
  • CBT and Problem-Solving intents tended to produce longer, more directive conversations, while Open-Ended reflection prompts sometimes yielded shorter exchanges; overall, dialog dynamics varied by the intent implementation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.