Skip to main content
QUICK REVIEW

[Paper Review] KANDINSKYPatterns -- An experimental exploration environment for Pattern Analysis and Machine Intelligence

Andreas Holzinger, Anna Saranti|arXiv (Cornell University)|Feb 28, 2021
Explainable Artificial Intelligence (XAI)72 references4 citations
TL;DR

This paper introduces KANDINSKYPatterns (KP), a novel experimental environment for pattern analysis and machine intelligence that enables controlled, human-interpretable visual concept learning. Inspired by Kandinsky’s theory of compositional perception and Hubel & Wiesel’s neural basis of vision, KP provides computationally controllable, ground-truth-verified patterns with hierarchical, compositional concepts—enabling benchmarking of generalization, reasoning, and explainability in AI systems.

ABSTRACT

Machine intelligence is very successful at standard recognition tasks when having high-quality training data. There is still a significant gap between machine-level pattern recognition and human-level concept learning. Humans can learn under uncertainty from only a few examples and generalize these concepts to solve new problems. The growing interest in explainable machine intelligence, requires experimental environments and diagnostic tests to analyze weaknesses in existing approaches to drive progress in the field. In this paper, we discuss existing diagnostic tests and test data sets such as CLEVR, CLEVERER, CLOSURE, CURI, Bongard-LOGO, V-PROM, and present our own experimental environment: The KANDINSKYPatterns, named after the Russian artist Wassily Kandinksy, who made theoretical contributions to compositivity, i.e. that all perceptions consist of geometrically elementary individual components. This was experimentally proven by Hubel &Wiesel in the 1960s and became the basis for machine learning approaches such as the Neocognitron and the even later Deep Learning. While KANDINSKYPatterns have computationally controllable properties on the one hand, bringing ground truth, they are also easily distinguishable by human observers, i.e., controlled patterns can be described by both humans and algorithms, making them another important contribution to international research in machine intelligence.

Motivation & Objective

  • To bridge the gap between machine-level pattern recognition and human-level concept learning by creating a testbed for evaluating generalization and reasoning under uncertainty.
  • To provide a diagnostic environment that supports both human and algorithmic understanding of visual concepts, enabling direct comparison of AI and human performance.
  • To foster research in explainable AI by offering a dataset with ground-truth annotations and compositional, hierarchical concepts that reflect real-world cognitive processes.
  • To support the development of new neural network architectures and hybrid symbolic-AI methods through a structured, extensible benchmark with clear linguistic and structural definitions.
  • To enable systematic evaluation of few-shot learning, compositional generalization, and counterfactual reasoning in vision models using a curated, evolving dataset.

Proposed method

  • Designing a visual dataset based on geometric primitives (shapes, colors, positions) inspired by Kandinsky’s theory of compositional perception and Hubel & Wiesel’s discovery of hierarchical visual processing in the brain.
  • Creating a set of controlled, computationally verifiable patterns with ground truth, ensuring both algorithmic and human interpretability.
  • Defining a hierarchical, compositional concept space that includes basic concepts (e.g., shape, color, count) and higher-order relations (e.g., arithmetic, spatial relations).
  • Implementing a grammar or domain-specific language to express concepts and their ambiguities, supporting probabilistic and compositional reasoning.
  • Generating test splits that systematically evaluate generalization, compositionality, and robustness across training and test distributions.
  • Enabling unconstrained data generation for developers and researchers while preserving alignment with medical and cognitive reasoning benchmarks.

Experimental results

Research questions

  • RQ1How well can state-of-the-art neural networks generalize to compositional and hierarchical visual concepts in a controlled, interpretable environment?
  • RQ2To what extent can AI systems learn and reason about abstract visual concepts with minimal supervision, mimicking human few-shot concept learning?
  • RQ3How do model predictions compare to human performance in concept recognition, description, and disentanglement on the same visual patterns?
  • RQ4Can explainable AI (xAI) methods effectively diagnose failures in model reasoning on compositional visual tasks?
  • RQ5How can the dataset be extended to support counterfactual reasoning and causal understanding in vision models?

Key findings

  • KANDINSKYPatterns provide a unique combination of computational control, ground truth, and human interpretability, enabling direct benchmarking of AI and human concept learning.
  • The dataset supports hierarchical and compositional concept learning, allowing for tasks involving arithmetic, spatial relations, and relational reasoning beyond simple classification.
  • Compared to existing benchmarks like CLEVR and Bongard-LOGO, KP emphasizes compositional structure and supports richer linguistic and symbolic expression of visual concepts.
  • The dataset is designed to support future extensions into real-world domains, such as medical imaging, by aligning concept grammar with clinical diagnostic language.
  • The framework enables systematic evaluation of generalization and robustness, particularly in few-shot and out-of-distribution settings.
  • The integration of xAI methods with KP allows for deeper diagnostic analysis of model behavior, revealing both successes and failure modes in reasoning and generalization.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.