[论文解读] Understanding Users' Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level
本研究通过107名用户的511次不满意的互动,采用混合方法分析了用户对ChatGPT回复的不满情绪。研究识别出七类不满,其中意图误解最为常见,而准确性问题最为严重;尽管用户采用了如意图具体化等策略,72%的不满仍未能解决,而这些策略既常见又有效。用户的知识水平显著影响不满模式与解决行为,低知识用户面临更多准确性问题,且使用策略更少。
Large language models (LLMs) with chat-based capabilities, such as ChatGPT, are widely used in various workflows. However, due to a limited understanding of these large-scale models, users struggle to use this technology and experience different kinds of dissatisfaction. Researchers have introduced several methods, such as prompt engineering, to improve model responses. However, they focus on enhancing the model's performance in specific tasks, and little has been investigated on how to deal with the user dissatisfaction resulting from the model's responses. Therefore, with ChatGPT as the case study, we examine users' dissatisfaction along with their strategies to address the dissatisfaction. After organizing users' dissatisfaction with LLM into seven categories based on a literature review, we collected 511 instances of dissatisfactory ChatGPT responses from 107 users and their detailed recollections of dissatisfactory experiences, which we released as a publicly accessible dataset. Our analysis reveals that users most frequently experience dissatisfaction when ChatGPT fails to grasp their intentions, while they rate the severity of dissatisfaction related to accuracy the highest. We also identified four tactics users employ to address their dissatisfaction and their effectiveness. We found that users often do not use any tactics to address their dissatisfaction, and even when using tactics, 72% of dissatisfaction remained unresolved. Moreover, we found that users with low knowledge of LLMs tend to face more dissatisfaction on accuracy while they often put minimal effort in addressing dissatisfaction. Based on these findings, we propose design implications for minimizing user dissatisfaction and enhancing the usability of chat-based LLM.
研究动机与目标
- 理解用户在与基于聊天的大型语言模型(如ChatGPT)互动时所经历的不满类型。
- 识别并评估用户在持续对话中用于解决不满的策略。
- 研究用户对大型语言模型的知识水平如何影响其不满及解决行为。
- 为改进基于聊天的大型语言模型服务的可用性与用户体验提供设计启示。
提出的方法
- 通过系统性文献回顾,基于大型语言模型响应失败的类型,将用户侧的不满划分为七类。
- 收集并分析了来自107名用户的511个真实世界中不满意的ChatGPT互动实例,包括详细的回忆记录与对话日志。
- 通过定性分析,将用户响应策略划分为四类:意图具体化、重述、澄清请求和提示复用。
- 使用用户自评分数测量不满严重性与策略有效性,并与实际对话日志进行验证。
- 根据知识水平(高 vs. 低)对用户进行分组,分析其在不满频率、严重性及解决行为上的差异。
- 公开发布了一个可供未来研究使用的用户报告的不满案例数据集。
实验结果
研究问题
- RQ1用户在与ChatGPT互动时,主要经历哪些类型的不满?
- RQ2用户采用哪些策略来解决其不满,这些策略的有效性如何?
- RQ3用户对大型语言模型的知识水平如何影响不满的频率、严重性及解决情况?
- RQ4在用户经验中,哪些不满类别最为频繁,哪些最为严重?
主要发现
- 最常见的不满类别是‘意图理解’失败,即ChatGPT未能理解用户的实际请求意图。
- 最严重的不满与‘信息准确性’问题相关,即回复中包含事实性错误或幻觉内容。
- 尽管使用了策略,仍有72%的不满实例未得到解决,表明当前用户策略存在显著局限。
- ‘意图具体化’策略在所有策略中使用频率最高,且效果最佳。
- 对大型语言模型知识水平较低的用户,其遭遇的与准确性及内容深度相关的不满显著更多,且更少使用任何解决策略。
- 低知识用户更可能遭遇拒绝回复(D_refuse),且常无修改地重复使用提示,表明缺乏有效的适应策略。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。