Skip to main content
QUICK REVIEW

[論文レビュー] Understanding Users' Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level

Yoonsu Kim, J. K.W. Lee|arXiv (Cornell University)|Nov 13, 2023
Artificial Intelligence in Healthcare and Education参考文献 88被引用数 4
ひとこと要約

本研究は、107名のユーザーから得られた511件の不満を示す対話の混合研究的手法による分析を通じて、ChatGPTの応答に対するユーザーの不満を調査している。7つの不満カテゴリを特定し、意図の誤解が最も頻度が高く、正確性の問題が最も深刻であることが判明した。一方で、ユーザーが意図の明確化などの戦略を用いても、72%の不満は解決されていないことが明らかになった。この戦略は一般的かつ効果的であるが、知識水準が不満のパターンと解決への取り組みに顕著な影響を与え、低知識ユーザーは正確性の問題にさらされやすく、戦略の使用頻度も低いことが分かった。

ABSTRACT

Large language models (LLMs) with chat-based capabilities, such as ChatGPT, are widely used in various workflows. However, due to a limited understanding of these large-scale models, users struggle to use this technology and experience different kinds of dissatisfaction. Researchers have introduced several methods, such as prompt engineering, to improve model responses. However, they focus on enhancing the model's performance in specific tasks, and little has been investigated on how to deal with the user dissatisfaction resulting from the model's responses. Therefore, with ChatGPT as the case study, we examine users' dissatisfaction along with their strategies to address the dissatisfaction. After organizing users' dissatisfaction with LLM into seven categories based on a literature review, we collected 511 instances of dissatisfactory ChatGPT responses from 107 users and their detailed recollections of dissatisfactory experiences, which we released as a publicly accessible dataset. Our analysis reveals that users most frequently experience dissatisfaction when ChatGPT fails to grasp their intentions, while they rate the severity of dissatisfaction related to accuracy the highest. We also identified four tactics users employ to address their dissatisfaction and their effectiveness. We found that users often do not use any tactics to address their dissatisfaction, and even when using tactics, 72% of dissatisfaction remained unresolved. Moreover, we found that users with low knowledge of LLMs tend to face more dissatisfaction on accuracy while they often put minimal effort in addressing dissatisfaction. Based on these findings, we propose design implications for minimizing user dissatisfaction and enhancing the usability of chat-based LLM.

研究の動機と目的

  • チャットベースのLLM(例:ChatGPT)との対話において、ユーザーが経験する不満の種類を理解すること。
  • 継続的な会話の中で不満を解消するためにユーザーが採用する戦略を特定・評価すること。
  • ユーザーのLLMに関する知識水準が、不満の頻度・深刻度および解決行動に与える影響を検討すること。
  • チャットベースのLLMサービスの使いやすさとユーザーエクスペリエンスを向上させるための設計的示唆を提供すること。

提案手法

  • LLM応答の失敗に基づき、ユーザー側の不満を7つの明確なタイプに分類するため、体系的文献レビューを実施した。
  • 107名のユーザーから得た511件の現実世界の不満を示すChatGPT対話の実例を収集・分析し、詳細な記憶と会話ログを含めた。
  • 定性的分析を通じて、ユーザーの応答戦略を4つのカテゴリに分類した:意図の明確化、言い換え、確認要求、プロンプトの再利用。
  • 自らの評価スコアを用いて不満の深刻度と戦略の有効性を測定し、実際の会話ログと照合して妥当性を検証した。
  • 知識水準(高・低)に応じてユーザーをセグメント化し、不満の頻度、深刻度、解決行動の差を分析した。
  • 今後の研究のために、ユーザーが報告した不満事例の公開可能なデータセットを公開した。

実験結果

リサーチクエスチョン

  • RQ1ユーザーがChatGPTの応答と対話する際に経験する主な不満の種類は何か?
  • RQ2ユーザーはどのような戦略を不満解消のために用いるのか?また、その有効性はいかほどか?
  • RQ3ユーザーのLLMに関する知識水準は、不満の頻度・深刻度および解決行動にどのように影響を与えるか?
  • RQ4ユーザー経験において、最も頻度が高く、最も深刻な不満カテゴリはどれか?

主な発見

  • 最も頻度の高かった不満カテゴリは「意図の理解不能」であり、ChatGPTがユーザーの意図した要求を把握できなかったことによる。
  • 最も深刻な不満は「情報の正確性」の問題に関連しており、応答に事実誤認や空想的記述(ホールーシュレーション)が含まれていた。
  • 戦略を用いても、72%の不満事例が解決されていないため、現在のユーザー戦略には顕著な限界があることが示された。
  • 「意図の明確化」は、最も頻繁に使用され、かつ最も効果的な戦略であった。
  • LLMに関する知識が低いユーザーは、正確性や内容の深さに関する不満を著しく多く経験し、解決戦略を用いる頻度も低かった。
  • 低知識ユーザーは拒否応答(D_refuse)にさらされやすく、しばしば変更なしに同じプロンプトを繰り返し使用しており、効果的な適応戦略が欠如していることが示唆された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。