[Paper Review] Understanding Users' Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level
This study investigates user dissatisfaction with ChatGPT responses through a mixed-methods analysis of 511 dissatisfying interactions from 107 users. It identifies seven dissatisfaction categories, with intent misunderstanding being most frequent and accuracy issues the most severe, while revealing that 72% of dissatisfaction remains unresolved despite users employing tactics like intent concretization, which is both common and effective. Knowledge level significantly affects dissatisfaction patterns and resolution efforts, with low-knowledge users facing more accuracy issues and using fewer tactics.
Large language models (LLMs) with chat-based capabilities, such as ChatGPT, are widely used in various workflows. However, due to a limited understanding of these large-scale models, users struggle to use this technology and experience different kinds of dissatisfaction. Researchers have introduced several methods, such as prompt engineering, to improve model responses. However, they focus on enhancing the model's performance in specific tasks, and little has been investigated on how to deal with the user dissatisfaction resulting from the model's responses. Therefore, with ChatGPT as the case study, we examine users' dissatisfaction along with their strategies to address the dissatisfaction. After organizing users' dissatisfaction with LLM into seven categories based on a literature review, we collected 511 instances of dissatisfactory ChatGPT responses from 107 users and their detailed recollections of dissatisfactory experiences, which we released as a publicly accessible dataset. Our analysis reveals that users most frequently experience dissatisfaction when ChatGPT fails to grasp their intentions, while they rate the severity of dissatisfaction related to accuracy the highest. We also identified four tactics users employ to address their dissatisfaction and their effectiveness. We found that users often do not use any tactics to address their dissatisfaction, and even when using tactics, 72% of dissatisfaction remained unresolved. Moreover, we found that users with low knowledge of LLMs tend to face more dissatisfaction on accuracy while they often put minimal effort in addressing dissatisfaction. Based on these findings, we propose design implications for minimizing user dissatisfaction and enhancing the usability of chat-based LLM.
Motivation & Objective
- To understand the types of dissatisfaction users experience when interacting with chat-based LLMs like ChatGPT.
- To identify and evaluate user strategies for resolving dissatisfaction during ongoing conversations.
- To examine how users' knowledge levels about LLMs influence their dissatisfaction and resolution behaviors.
- To provide design implications for improving the usability and user experience of chat-based LLM services.
Proposed method
- Conducted a systematic literature review to categorize user-side dissatisfaction into seven distinct types based on LLM response failures.
- Collected and analyzed 511 real-world instances of dissatisfying ChatGPT interactions from 107 users, including detailed recollections and conversation logs.
- Classified user response tactics into four categories through qualitative analysis: intent concretization, rephrasing, clarification requests, and prompt reuse.
- Measured dissatisfaction severity and tactic effectiveness using self-reported user scores, validated against actual conversation logs.
- Segmented users by knowledge level (high vs. low) to analyze differences in dissatisfaction frequency, severity, and resolution behavior.
- Released a publicly accessible dataset of user-reported dissatisfaction cases for future research.
Experimental results
Research questions
- RQ1What are the primary types of dissatisfaction users experience when interacting with ChatGPT responses?
- RQ2What tactics do users employ to resolve their dissatisfaction, and how effective are these tactics?
- RQ3How does the user’s knowledge level about LLMs influence the frequency, severity, and resolution of dissatisfaction?
- RQ4Which dissatisfaction categories are most frequent and most severe in user experiences?
Key findings
- The most frequent dissatisfaction category was 'intent understanding' failure, where ChatGPT failed to grasp the user’s intended request.
- The most severe dissatisfaction was associated with 'information accuracy' issues, where responses contained factual errors or hallucinations.
- Despite using tactics, 72% of dissatisfaction instances remained unresolved, indicating significant limitations in current user strategies.
- The tactic 'intent concretization' was both the most frequently used and the most effective in resolving dissatisfaction.
- Users with low knowledge about LLMs experienced significantly more dissatisfaction related to accuracy and content depth, and were less likely to use any resolution tactics.
- Low-knowledge users were more likely to encounter refusal responses (D_refuse) and often reused prompts without modification, suggesting a lack of effective adaptation strategies.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.