Skip to main content
QUICK REVIEW

[Paper Review] How Good is ChatGPT in Giving Advice on Your Visualization Design?

Nam Wook Kim, Ahn, Yongsu|arXiv (Cornell University)|Oct 14, 2023
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This study evaluates ChatGPT's ability to provide data visualization design advice by comparing its responses to human expert feedback on real practitioner questions. Using a mixed-method approach—analyzing forum posts and conducting a user study—results show ChatGPT matches or exceeds human responses in breadth, clarity, and coverage, though humans are preferred for depth, nuance, and conversational adaptability.

ABSTRACT

Data visualization creators often lack formal training, resulting in a knowledge gap in design practice. Large language models such as ChatGPT, with their vast internet-scale training data, offer transformative potential to address this gap. In this study, we used both qualitative and quantitative methods to investigate how well ChatGPT can address visualization design questions. First, we quantitatively compared the ChatGPT-generated responses with anonymous online Human replies to data visualization questions on the VisGuides user forum. Next, we conducted a qualitative user study examining the reactions and attitudes of practitioners toward ChatGPT as a visualization design assistant. Participants were asked to bring their visualizations and design questions and received feedback from both Human experts and ChatGPT in randomized order. Our findings from both studies underscore ChatGPT's strengths, particularly its ability to rapidly generate diverse design options, while also highlighting areas for improvement, such as nuanced contextual understanding and fluid interaction dynamics beyond the chat interface. Drawing on these insights, we discuss design considerations for future LLM-based design feedback systems.

Motivation & Objective

  • To address the knowledge gap in data visualization design among non-expert practitioners who lack formal training.
  • To investigate whether large language models like ChatGPT can serve as viable alternatives or complements to human experts in providing design feedback.
  • To understand practitioners’ perceptions, preferences, and experiences when receiving feedback from both ChatGPT and human experts.
  • To establish a foundation for benchmarking LLMs in data visualization by identifying key evaluation metrics and real-world use cases.
  • To explore the potential integration of LLM-based assistants into visualization tools and educational settings.

Proposed method

  • Conducted a quantitative analysis of 100+ questions from the VisGuide forum, comparing ChatGPT-generated responses to human expert replies using six evaluation metrics: coverage, topicality, breadth, clarity, depth, and actionability.
  • Performed a mixed-method user study with 21 practitioners who brought their own visualizations and received feedback from both ChatGPT and human experts in randomized order.
  • Collected structured experience surveys and conducted post-interviews to assess user perceptions, preferences, and feedback quality.
  • Used thematic analysis to identify patterns in user feedback, focusing on communication dynamics, perceived expertise, and trust in AI vs. human advice.
  • Evaluated ChatGPT’s performance across diverse visualization design challenges, including chart selection, data encoding, and visual clarity.
  • Proposed a framework for future benchmarking of LLMs in data visualization using the collected dataset and evaluation metrics.
Figure 1: Methodology Overview: The methodology comprises two key phases. In the first phase, questions answered within a forum space by Human respondents are explored and then presented to ChatGPT . The second phase involves a feedback session in which users solicit visualization design feedback fr
Figure 1: Methodology Overview: The methodology comprises two key phases. In the first phase, questions answered within a forum space by Human respondents are explored and then presented to ChatGPT . The second phase involves a feedback session in which users solicit visualization design feedback fr

Experimental results

Research questions

  • RQ1RQ1: Can ChatGPT rival human expertise in data visualization knowledge, particularly in terms of response quality across key metrics?
  • RQ2RQ2: How do visualization practitioners perceive and react to feedback from ChatGPT compared to feedback from human experts?
  • RQ3RQ3: What are the strengths and limitations of ChatGPT in providing actionable, nuanced, and context-aware design advice?
  • RQ4RQ4: How can LLM-based assistants be effectively integrated into visualization workflows and educational contexts?

Key findings

  • ChatGPT outperformed human responses in coverage, breadth, and clarity, providing comprehensive and well-structured answers to visualization design questions.
  • Human experts were consistently preferred by practitioners due to their ability to engage in fluid, adaptive conversations and offer deeper, context-sensitive critiques.
  • ChatGPT demonstrated strong performance in generating diverse design suggestions and ideation, particularly useful in early-stage design exploration.
  • Despite its broad knowledge base, ChatGPT struggled with depth and critical thinking, often producing generic or surface-level feedback.
  • Participants reported that ChatGPT’s passive, affirmative communication style limited its ability to challenge assumptions or guide users through complex design decisions.
  • The study identified a need for future LLMs to improve visual perception capabilities and adopt more proactive, dialogic interaction patterns to better support design reasoning.
Figure 2: Comparison of metric scores between Human and ChatGPT responses: the chart showcases the existence of significant differences for the metrics of breadth, clarity, and coverage, while the scores for actionablity, depth and topicality were comparable for ChatGPT and Human response ratings. E
Figure 2: Comparison of metric scores between Human and ChatGPT responses: the chart showcases the existence of significant differences for the metrics of breadth, clarity, and coverage, while the scores for actionablity, depth and topicality were comparable for ChatGPT and Human response ratings. E

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.