[Paper Review] SEVA: Leveraging sketches to evaluate alignment between human and machine visual abstraction
SEVA introduces a benchmark dataset of 90K human-generated sketches across 128 object concepts with varying sparsity due to time constraints, enabling evaluation of how closely machine vision models align with human visual abstraction. The study reveals that while top models predict human sketch recognition performance and uncertainty patterns better than baseline models, a significant gap remains in full behavioral alignment.
Sketching is a powerful tool for creating abstract images that are sparse but meaningful. Sketch understanding poses fundamental challenges for general-purpose vision algorithms because it requires robustness to the sparsity of sketches relative to natural visual inputs and because it demands tolerance for semantic ambiguity, as sketches can reliably evoke multiple meanings. While current vision algorithms have achieved high performance on a variety of visual tasks, it remains unclear to what extent they understand sketches in a human-like way. Here we introduce SEVA, a new benchmark dataset containing approximately 90K human-generated sketches of 128 object concepts produced under different time constraints, and thus systematically varying in sparsity. We evaluated a suite of state-of-the-art vision algorithms on their ability to correctly identify the target concept depicted in these sketches and to generate responses that are strongly aligned with human response patterns on the same sketch recognition task. We found that vision algorithms that better predicted human sketch recognition performance also better approximated human uncertainty about sketch meaning, but there remains a sizable gap between model and human response patterns. To explore the potential of models that emulate human visual abstraction in generative tasks, we conducted further evaluations of a recently developed sketch generation algorithm (Vinker et al., 2022) capable of generating sketches that vary in sparsity. We hope that public release of this dataset and evaluation protocol will catalyze progress towards algorithms with enhanced capacities for human-like visual abstraction.
Motivation & Objective
- To develop a standardized benchmark for evaluating how well machine vision models understand freehand sketches in a way that aligns with human visual abstraction.
- To investigate the extent to which state-of-the-art vision models replicate human response patterns in sketch recognition, including sensitivity to sketch sparsity and semantic ambiguity.
- To explore the potential of models that emulate human visual abstraction in generative tasks, particularly sketch generation with controlled sparsity.
- To provide a public dataset and evaluation protocol that supports progress toward unified computational theories of human-like visual abstraction.
- To identify gaps between model and human performance in sketch understanding, especially regarding uncertainty and contextual interpretation.
Proposed method
- Collecting approximately 90,000 human-generated sketches of 128 object concepts using a web-based digital interface under varying time constraints to induce systematic variation in sketch sparsity.
- Recording detailed stroke-level dynamics (e.g., timing, sequence, pressure) during sketch production to capture moment-to-moment decision-making processes.
- Evaluating state-of-the-art vision models on sketch recognition tasks using the dataset to measure alignment with human response patterns.
- Measuring model performance not only on accuracy but also on their ability to predict human uncertainty and response variability across different sketch abstractions.
- Conducting additional evaluations on a sketch generation model (CLIPasso) to assess its capacity to produce sketches that match human-like abstraction patterns.
- Establishing a public benchmark with standardized evaluation protocols to enable reproducible comparison across models and future research.
Experimental results
Research questions
- RQ1To what extent do state-of-the-art vision models replicate human response patterns in recognizing sketches of varying sparsity?
- RQ2How well do vision models predict human uncertainty about the meaning of ambiguous or sparse sketches compared to human behavior?
- RQ3Can generative models produce sketches that align with human visual abstraction patterns in terms of sparsity and interpretability?
- RQ4What is the performance gap between the best-performing models and human consistency in sketch recognition across diverse abstraction levels?
- RQ5How do input modality (e.g., mouse vs. stylus), cultural background, and artistic training affect sketch production and recognition behavior, and how can models account for such variation?
Key findings
- Vision models that better predicted human sketch recognition performance also showed stronger alignment with human uncertainty patterns, indicating improved modeling of semantic ambiguity.
- Despite high accuracy on sketch recognition, even the best-performing models fell short of human consistency baselines, particularly in handling sparse and ambiguous sketches.
- The dataset revealed systematic variation in sketch sparsity due to time constraints, with clearer distinctions in human recognition patterns across different abstraction levels.
- Models trained on natural images still lack full representational alignment with human visual abstraction, especially when generalizing to sketch distributions.
- The CLIPasso sketch generation model demonstrated potential for producing human-like abstractions, but its output still diverged from human sketching behavior in key aspects.
- The study highlights a persistent gap in behavioral alignment between models and humans, particularly in how uncertainty and context are handled during sketch interpretation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.