[Paper Review] A Survey of Visual Analytics Techniques for Machine Learning
This paper presents a comprehensive survey of visual analytics techniques for machine learning, organizing 259 studies into three stages: before, during, and after model building. It introduces a structured taxonomy, highlights key tasks such as data quality improvement, model diagnosis, and concept drift analysis, and identifies six future research directions to enhance explainability, trust, and usability in machine learning pipelines.
Visual analytics for machine learning has recently evolved as one of the most exciting areas in the field of visualization. To better identify which research topics are promising and to learn how to apply relevant techniques in visual analytics, we systematically review 259 papers published in the last ten years together with representative works before 2010. We build a taxonomy, which includes three first-level categories: techniques before model building, techniques during model building, and techniques after model building. Each category is further characterized by representative analysis tasks, and each task is exemplified by a set of recent influential works. We also discuss and highlight research challenges and promising potential future research opportunities useful for visual analytics researchers.
Motivation & Objective
- To provide a systematic, comprehensive review of visual analytics techniques for machine learning across the entire machine learning pipeline.
- To identify and categorize key research trends and representative works in visual analytics for machine learning over the past decade.
- To highlight underexplored research directions and open challenges in improving model explainability, data quality, and system usability.
- To support researchers and practitioners in understanding how visualization can enhance trust, reliability, and interpretability in machine learning workflows.
- To establish a taxonomy based on the machine learning lifecycle, covering pre-, in-, and post-modeling phases with task-specific examples.
Proposed method
- Conducted a systematic manual review of 259 papers from top-tier visualization venues (e.g., InfoVis, VAST, IEEE TVCG, SciVis) published between 2010 and 2020.
- Classified visual analytics techniques into three stages: before model building (data and feature quality), during model building (model diagnosis and steering), and after model building (model and data understanding).
- Abstracted representative analysis tasks within each stage, such as data quality assessment, feature selection, model interpretation, and concept drift detection.
- Selected influential representative works to exemplify each task, emphasizing interactive and explainable visualization techniques.
- Integrated insights from prior surveys and ontologies to refine the taxonomy and ensure coverage of emerging trends like multi-modal learning and online concept drift.
- Identified six key future research directions through synthesis of recurring challenges and gaps in current literature.
Experimental results
Research questions
- RQ1How can visual analytics techniques improve data and feature quality before machine learning model training?
- RQ2What visualization strategies support real-time diagnosis, interpretation, and interactive refinement of machine learning models during training?
- RQ3How can visual analytics facilitate the understanding of model behavior and data relationships after deployment, especially in complex scenarios like multi-modal data?
- RQ4What are the key challenges in detecting and analyzing concept drift in streaming machine learning applications, and how can visualization support human-in-the-loop diagnosis?
- RQ5What are the most promising future research directions for integrating visual analytics with machine learning to improve explainability and reliability?
Key findings
- The survey identifies 259 relevant papers from top-tier visualization venues, forming a comprehensive foundation for understanding visual analytics in machine learning.
- A three-stage taxonomy—before, during, and after model building—effectively organizes visual analytics techniques across the machine learning lifecycle.
- Key pre-modeling tasks include data quality assessment, missing data handling, and feature selection, with visualization enabling intuitive data inspection and labeling correction.
- During model building, visual analytics supports model diagnosis, hyperparameter tuning, and interpretation of complex models like deep neural networks through interactive visual interfaces.
- Post-modeling analysis includes understanding model predictions, detecting data drift, and interpreting multi-modal data relationships, with visualization enabling comparative analysis of data distributions over time.
- Six major future research directions are identified: improving data quality in weakly supervised learning, enabling online model diagnosis, supporting intelligent model refinement, enhancing multi-modal data understanding, enabling concept drift analysis with human-in-the-loop, and developing unified visualizations for joint multi-modal representations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.