Skip to main content
QUICK REVIEW

[论文解读] How Good is ChatGPT in Giving Advice on Your Visualization Design?

Nam Wook Kim, Ahn, Yongsu|arXiv (Cornell University)|Oct 14, 2023
Artificial Intelligence in Healthcare and EducationMedicine被引用 3
一句话总结

本研究通过将ChatGPT对实际从业者问题的响应与人类专家反馈进行对比,评估其在提供数据可视化设计建议方面的能力。采用混合方法研究——分析论坛帖子并开展用户研究——结果表明,ChatGPT在广度、清晰度和覆盖范围方面表现与人类相当甚至更优,尽管在深度、细微差别和对话适应性方面人类仍更受青睐。

ABSTRACT

Data visualization creators often lack formal training, resulting in a knowledge gap in design practice. Large language models such as ChatGPT, with their vast internet-scale training data, offer transformative potential to address this gap. In this study, we used both qualitative and quantitative methods to investigate how well ChatGPT can address visualization design questions. First, we quantitatively compared the ChatGPT-generated responses with anonymous online Human replies to data visualization questions on the VisGuides user forum. Next, we conducted a qualitative user study examining the reactions and attitudes of practitioners toward ChatGPT as a visualization design assistant. Participants were asked to bring their visualizations and design questions and received feedback from both Human experts and ChatGPT in randomized order. Our findings from both studies underscore ChatGPT's strengths, particularly its ability to rapidly generate diverse design options, while also highlighting areas for improvement, such as nuanced contextual understanding and fluid interaction dynamics beyond the chat interface. Drawing on these insights, we discuss design considerations for future LLM-based design feedback systems.

研究动机与目标

  • 解决非专家从业者在数据可视化设计方面存在的知识缺口,这些从业者缺乏正式培训。
  • 探究大型语言模型(如ChatGPT)是否可作为人类专家在提供设计反馈方面的可行替代品或补充。
  • 了解从业者在接收来自ChatGPT和人类专家的反馈时的感知、偏好和体验。
  • 通过识别关键评估指标和真实应用场景,为大语言模型在数据可视化领域的基准测试奠定基础。
  • 探索基于大语言模型的助手在可视化工具和教育场景中的潜在整合方式。

提出的方法

  • 对VisGuide论坛中100多个问题进行了定量分析,使用六项评估指标(覆盖度、相关性、广度、清晰度、深度和可操作性)对比ChatGPT生成的回应与人类专家的回复。
  • 开展了一项混合方法用户研究,21名从业者带来了自己的可视化作品,以随机顺序从ChatGPT和人类专家处获得反馈。
  • 收集结构化体验调查并进行访谈,以评估用户感知、偏好和反馈质量。
  • 使用主题分析法识别用户反馈中的模式,重点关注沟通动态、对专业能力的感知以及对AI与人类建议的信任度。
  • 在多种可视化设计挑战中评估ChatGPT的表现,包括图表选择、数据编码和视觉清晰度。
  • 基于收集的数据集和评估指标,提出一个用于未来大语言模型在数据可视化领域基准测试的框架。
Figure 1: Methodology Overview: The methodology comprises two key phases. In the first phase, questions answered within a forum space by Human respondents are explored and then presented to ChatGPT . The second phase involves a feedback session in which users solicit visualization design feedback fr
Figure 1: Methodology Overview: The methodology comprises two key phases. In the first phase, questions answered within a forum space by Human respondents are explored and then presented to ChatGPT . The second phase involves a feedback session in which users solicit visualization design feedback fr

实验结果

研究问题

  • RQ1RQ1:ChatGPT能否在关键指标上媲美人类专家在数据可视化知识方面的专业水平,特别是在回应质量方面?
  • RQ2RQ2:可视化从业者如何感知并回应来自ChatGPT的反馈,与人类专家的反馈相比有何差异?
  • RQ3RQ3:ChatGPT在提供可操作性、细致入微且具备上下文感知的设计建议方面,其优势与局限是什么?
  • RQ4RQ4:如何有效将基于大语言模型的助手整合到可视化工作流和教育场景中?

主要发现

  • ChatGPT在覆盖度、广度和清晰度方面优于人类回应,能够为可视化设计问题提供全面且结构良好的答案。
  • 从业者始终更偏好人类专家,因其具备流畅、灵活的对话能力,并能提供更深入、上下文敏感的批判性意见。
  • ChatGPT在生成多样化的设计建议和创意构思方面表现强劲,尤其适用于早期阶段的设计探索。
  • 尽管知识库广泛,ChatGPT在深度和批判性思维方面仍显不足,常产生通用或表面化的反馈。
  • 参与者报告称,ChatGPT被动且肯定的沟通风格限制了其挑战假设或引导用户进行复杂设计决策的能力。
  • 本研究识别出未来大语言模型需提升视觉感知能力,并采用更具主动性、对话性的互动模式,以更好地支持设计推理。
Figure 2: Comparison of metric scores between Human and ChatGPT responses: the chart showcases the existence of significant differences for the metrics of breadth, clarity, and coverage, while the scores for actionablity, depth and topicality were comparable for ChatGPT and Human response ratings. E
Figure 2: Comparison of metric scores between Human and ChatGPT responses: the chart showcases the existence of significant differences for the metrics of breadth, clarity, and coverage, while the scores for actionablity, depth and topicality were comparable for ChatGPT and Human response ratings. E

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。