Skip to main content
QUICK REVIEW

[Paper Review] Categorization Axioms for Clustering Results

Yu Jian, Zongben Xu|arXiv (Cornell University)|Mar 9, 2014
Advanced Clustering Algorithms Research40 references3 citations
TL;DR

This paper proposes a formal axiomatic framework for clustering results based on categorization principles, introducing separation axioms that classify clustering outcomes as proper, overlapping, or improper. It establishes three design principles for clustering algorithms and validity indices—categorization equivalence, compactness, and separation—demonstrating consistency with major algorithms and indices like Xie-Beni and CVNN.

ABSTRACT

Cluster analysis has attracted more and more attention in the field of machine learning and data mining. Numerous clustering algorithms have been proposed and are being developed due to diverse theories and various requirements of emerging applications. Therefore, it is very worth establishing an unified axiomatic framework for data clustering. In the literature, it is an open problem and has been proved very challenging. In this paper, clustering results are axiomatized by assuming that an proper clustering result should satisfy categorization axioms. The proposed axioms not only introduce classification of clustering results and inequalities of clustering results, but also are consistent with prototype theory and exemplar theory of categorization models in cognitive science. Moreover, the proposed axioms lead to three principles of designing clustering algorithm and cluster validity index, which follow many popular clustering algorithms and cluster validity indices.

Motivation & Objective

  • To establish a unified axiomatic framework for clustering results, addressing the long-standing challenge of formalizing clustering in data analysis.
  • To define what constitutes a proper clustering result by formalizing categorization axioms that prevent trivial or degenerate outcomes.
  • To bridge clustering theory with cognitive science by aligning axioms with prototype and exemplar theories of categorization.
  • To derive general principles for designing clustering algorithms and cluster validity indices based on theoretical axioms.
  • To demonstrate that widely used clustering methods and indices inherently satisfy or align with the proposed axioms.

Proposed method

  • Introduces three core axioms: categorization equivalency, sample separation, and cluster separation, formalizing the structure of clustering results.
  • Defines improper clustering outcomes such as coincident, totally coincident, and uninformative partitions to exclude degenerate solutions.
  • Derives inequalities from the sample separation axiom to quantify the validity of clustering structures.
  • Proposes three algorithmic design principles: categorization equivalence, cluster compactness, and cluster separation, grounded in the axioms.
  • Applies the axioms to analyze and justify existing clustering algorithms and validity indices, such as fuzzy C-means and Xie-Beni Index.
  • Uses cognitive science theories (prototype and exemplar) to validate the theoretical consistency of the proposed axioms.

Experimental results

Research questions

  • RQ1What formal axioms can characterize a proper clustering result, excluding trivial or degenerate outcomes?
  • RQ2How can clustering results be systematically classified into proper, overlapping, or improper categories based on structural properties?
  • RQ3To what extent do established clustering algorithms and validity indices satisfy the proposed axioms and principles?
  • RQ4Can the axioms be used to derive general design principles for new clustering algorithms and validity indices?
  • RQ5How do the proposed axioms relate to cognitive categorization theories in psychology and cognitive science?

Key findings

  • The proposed categorization axioms successfully classify clustering results into proper, overlapping, and improper categories, with improper results including coincident and uninformative partitions.
  • The sample separation axiom leads to formal inequalities that quantify the structural validity of clustering outcomes.
  • The Xie-Beni index is shown to be inherently consistent with the proposed cluster separation and compactness principles, as it penalizes coincident clusters via its denominator.
  • Many popular clustering algorithms, including fuzzy C-means and model-based clustering, implicitly satisfy the proposed compactness and separation principles.
  • The cluster validity index CVNN is found to align with the proposed principles by jointly optimizing compactness and separation.
  • The axiomatic framework provides a theoretical foundation for designing new clustering algorithms and validity indices that avoid degenerate solutions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.