Skip to main content
QUICK REVIEW

[论文解读] Leveraging Large Language Models to Power Chatbots for Collecting User Self-Reported Data

Jing Wei, Sungdong Kim|arXiv (Cornell University)|Jan 14, 2023
AI in Service Interactions参考文献 85被引用 15
一句话总结

本论文研究如何为基于 GPT-3 的聊天机器人进行零-shot 提示,以收集自我报告的健康数据,并评估提示结构和个性提示如何影响数据收集与对话风格。

ABSTRACT

Large language models (LLMs) provide a new way to build chatbots by accepting natural language prompts. Yet, it is unclear how to design prompts to power chatbots to carry on naturalistic conversations while pursuing a given goal, such as collecting self-report data from users. We explore what design factors of prompts can help steer chatbots to talk naturally and collect data reliably. To this aim, we formulated four prompt designs with different structures and personas. Through an online study (N = 48) where participants conversed with chatbots driven by different designs of prompts, we assessed how prompt designs and conversation topics affected the conversation flows and users' perceptions of chatbots. Our chatbots covered 79% of the desired information slots during conversations, and the designs of prompts and topics significantly influenced the conversation flows and the data collection performance. We discuss the opportunities and challenges of building chatbots with LLMs.

研究动机与目标

  • 探索提示设计因素如何影响 LLM 驱动的聊天机器人在健康主题中收集用户自报数据。
  • 评估在不进行微调的情况下,GPT-3 驱动的聊天机器人之槽位填充性能与对话风格。
  • 确定用于与 LLM 基础聊天机器人进行健壮、自然互动和数据收集的设计准则。

提出的方法

  • 实现 16 个 GPT-3 驱动的聊天机器人(4 个主题 × 2 种格式 × 2 种个性)并使用结构化与描述性信息槽位提示。
  • 使用 davinci-text-002,统一的生成参数(温度 0.9,出现惩罚 0.6,频率惩罚 0.5)。
  • 提供基于网页的聊天界面,并在四个健康主题上进行一项涉及 48 名参与者的跨组在线研究。
  • 通过对是否获得预定义信息槽位进行人工编码来评估槽位填充。
  • 对对话行为进行编码以刻画对话流程和聊天机器人行为。
  • 分析参与者的退出调查,以评估对聊天机器人同理心和理解力的感知。
Figure 1 . An overview of our chatbot running on a large language model through zero-shot response generation, with only a prefix \footnotesize{A}⃝ consisting of persona modifier and information format and the ongoing dialogue history \footnotesize{B}⃝ . The example conversation is carried on about
Figure 1 . An overview of our chatbot running on a large language model through zero-shot response generation, with only a prefix \footnotesize{A}⃝ consisting of persona modifier and information format and the ongoing dialogue history \footnotesize{B}⃝ . The example conversation is carried on about

实验结果

研究问题

  • RQ1信息格式和人格修饰提示如何影响基于 LLM 的聊天机器人在槽位填充方面的表现?
  • RQ2GPT-3 驱动的聊天机器人在自报数据收集对话中是否能保持上下文并表现出同理心?
  • RQ3提示设计对任务导向的聊天机器人互动中的对话流程与用户感知有何影响?

主要发现

  • 在零-shot 提示的对话中,聊天机器人覆盖了所需信息槽的 79%。
  • 提示设计因素(信息格式与人格修饰符)显著影响对话流程和数据收集性能。
  • 聊天机器人通常表现出同理心,参与者有时感知响应的高准确性和有帮助性。
  • 研究证明了在不进行微调的情况下,基于 LLM 的聊天机器人用于收集自报数据并维持上下文与状态跟踪的可行性。
  • 对两个提示设计因素的系统性检验为使用 LLM 构建领域特定的零-shot 聊天机器人提供了见解。
Figure 2 . Prompt design combining two factors, information format and personality modifier, in the Food intake topic.
Figure 2 . Prompt design combining two factors, information format and personality modifier, in the Food intake topic.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。