[Paper Review] Leveraging Large Language Models to Power Chatbots for Collecting User Self-Reported Data
This paper studies how zero-shot prompts for GPT-3-powered chatbots can collect self-reported health data, evaluating how prompt structure and personality cues affect data collection and conversational style.
Large language models (LLMs) provide a new way to build chatbots by accepting natural language prompts. Yet, it is unclear how to design prompts to power chatbots to carry on naturalistic conversations while pursuing a given goal, such as collecting self-report data from users. We explore what design factors of prompts can help steer chatbots to talk naturally and collect data reliably. To this aim, we formulated four prompt designs with different structures and personas. Through an online study (N = 48) where participants conversed with chatbots driven by different designs of prompts, we assessed how prompt designs and conversation topics affected the conversation flows and users' perceptions of chatbots. Our chatbots covered 79% of the desired information slots during conversations, and the designs of prompts and topics significantly influenced the conversation flows and the data collection performance. We discuss the opportunities and challenges of building chatbots with LLMs.
Motivation & Objective
- Explore how prompt design factors influence LLM-driven chatbots for collecting user self-reported data across health topics.
- Evaluate slot-filling performance and conversation styles of chatbots powered by GPT-3 without fine-tuning.
- Identify design guidelines for robust, natural interactions and data collection with LLM-based chatbots.
Proposed method
- Implement 16 GPT-3 powered chatbots (4 topics × 2 formats × 2 personalities) with structured and descriptive information-slot prompts.
- Use davinci-text-002 with uniform generative parameters (temperature 0.9, presence penalty 0.6, frequency penalty 0.5).
- Provide a web-based chat interface and run a between-subject online study (N=48) across four health topics.
- Assess slot-filling by manual coding of whether predefined information slots were obtained.
- Code dialogue acts to characterize conversation flows and chatbot behaviors.
- Analyze participant exit surveys to gauge perceived chatbot empathy and understanding.

Experimental results
Research questions
- RQ1How do information-format and personality-modifier prompts influence slot-filling performance of LLM-driven chatbots?
- RQ2Do GPT-3 powered chatbots maintain context and demonstrate empathy in self-report data collection conversations?
- RQ3What is the impact of prompt design on conversation flow and user perception in task-oriented chatbot interactions?
Key findings
- Chatbots covered 79% of the desired information slots across dialogues with zero-shot prompts.
- Prompt design factors (information format and personality modifier) significantly influenced conversation flows and data collection performance.
- Chatbots generally responded empathetically, with participants sometimes perceiving high accuracy and helpfulness in responses.
- The study demonstrates feasibility of LLM-powered chatbots for collecting self-reports and maintaining context and state tracking without fine-tuning.
- A systematic examination of two prompt-design factors provides insights for scaffolding domain-specific, zero-shot chatbots using LLMs.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.