[Paper Review] Data Analysis in the Era of Generative AI
This paper explores how generative AI can transform data analysis by enabling non-experts to derive insights through natural language and multimodal interactions, reducing reliance on coding and specialized tools. It proposes AI-assisted workflows that translate high-level user intent into executable code, visualizations, and insights, while emphasizing human-centered design to enhance trust, usability, and collaboration across tools.
This paper explores the potential of AI-powered tools to reshape data analysis, focusing on design considerations and challenges. We explore how the emergence of large language and multimodal models offers new opportunities to enhance various stages of data analysis workflow by translating high-level user intentions into executable code, charts, and insights. We then examine human-centered design principles that facilitate intuitive interactions, build user trust, and streamline the AI-assisted analysis workflow across multiple apps. Finally, we discuss the research challenges that impede the development of these AI-based systems such as enhancing model capabilities, evaluating and benchmarking, and understanding end-user needs.
Motivation & Objective
- To investigate how generative AI can lower barriers to data analysis for non-expert users by translating natural language queries into actionable insights.
- To identify design principles that enhance user trust, streamline multi-tool workflows, and support iterative, exploratory data analysis.
- To address the gap in understanding user preferences for AI interaction modalities, autonomy levels, and personalization in AI-assisted data analysis.
- To highlight critical research challenges in model robustness, evaluation benchmarks, and infrastructure for high-quality, accessible data tables.
- To advocate for comprehensive user studies to guide the development of AI tools tailored to diverse user expertise and evolving analytical needs.
Proposed method
- Leveraging large language and multimodal models (e.g., GPT-4, GPT-4o, Phi-3) to interpret natural language and visual inputs and generate executable code, charts, and insights.
- Designing interactive, multi-modal interfaces that combine natural language, visual widgets, and code to support iterative, user-in-the-loop analysis workflows.
- Integrating AI assistance across multiple applications (e.g., data cleaning, visualization, reporting) to reduce tool-switching and maintain context continuity.
- Applying human-centered design principles to optimize user interaction, including modality selection (GUI + NL), AI autonomy levels, and feedback mechanisms.
- Proposing the use of multi-step, multi-modal data analysis scenarios as benchmarks to drive improvements in model planning, reasoning, and robustness.
- Developing infrastructure for indexing, ranking, and retrieving high-quality data tables from the web and enterprise sources to support AI-driven analysis suggestions.

Experimental results
Research questions
- RQ1How can generative AI systems effectively translate high-level user intentions into executable data analysis steps, including code and visualizations?
- RQ2What interaction modalities (e.g., GUI + natural language vs. chat-based) best support user understanding, control, and trust in AI-assisted data analysis?
- RQ3How can AI tools adapt to user expertise levels and evolving analytical goals without increasing cognitive load?
- RQ4What design patterns and interface components (e.g., AI-generated widgets) best support user control and verification of AI outputs across multiple applications?
- RQ5How can data infrastructure be enhanced to provide reliable, up-to-date, and privacy-compliant data tables for AI-driven analysis suggestions?
Key findings
- Users prefer hybrid interaction models combining GUI and natural language over pure chat-based assistants, as they provide better intent specification and constraint.
- AI-generated interactive widgets (e.g., for filtering or parameter adjustment) are more effective than direct AI execution, as they allow user control and reduce cognitive load.
- User reactions to AI suggestions vary significantly by statistical expertise: some find them helpful, while others perceive them as too basic or hard to interpret.
- Users require both procedural artifacts (e.g., code, explanations) and data artifacts (e.g., intermediate tables, visualizations) to verify and understand AI-generated analysis steps.
- Users value transparency and control over the context given to LLMs and prefer 'linter-like' assistants that flag inappropriate or incorrect AI usage.
- There remains a critical need for dynamic, multi-application AI coordination and personalized UIs that adapt to user history without frequent, disruptive changes.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.