[Paper Review] Taxonomic Provenance: Two Influential Primate Classifications Logically Aligned
This paper presents a logic-based framework to resolve taxonomic provenance between two editions of the Mammal Species of the World (MSW2 and MSW3) for primates, using Answer Set Programming to infer consistent alignments across 1,000+ taxonomic concepts. It identifies that approximately one in three name usages across editions lacks semantic congruence, demonstrating the feasibility of automated, machine-processable taxonomic provenance tracking.
Classification standards such as the Mammal Species of the World (MSW) aim to unify name usages at the global scale, but may nevertheless experience significant levels of taxonomic change from one edition to the next. This circumstance challenges the biodiversity and phylogenetic data communities to develop more granular identifiers to track taxonomic congruence and incongruence in ways that both humans and machines can process, i.e., to logically represent taxonomic provenance across multiple classification hierarchies. Here we show that reasoning over taxonomic provenance is feasible for two classifications of primates corresponding to the second and third MSW editions. Our approach entails three main components: (1) individuation of name usages as taxonomic concepts, (2) articulation of concepts via human-asserted Region Connection Calculus (RCC-5) relationships, and (3) the use of an Answer Set Programming toolkit to infer and visualize logically consistent alignments of these taxonomic input constraints. Our use case entails the Primates sec. Groves (1993; MSW2 - 317 taxonomic concepts; 233 at the species level) and Primates sec. Groves (2005; MSW3 - 483 taxonomic concepts; 376 at the species level). Using 402 concept-to-concept input articulations, the reasoning process yields a single, consistent alignment, and infers 153,111 Maximally Informative Relations that constitute a comprehensive provenance resolution map for every concept pair in the Primates sec. MSW2/MSW3. The entire alignment and various partitions facilitate quantitative analyses of name/meaning dissociation, revealing that approximately one in three paired name usages across treatments is not reliable - in the sense of the same name identifying congruent taxonomic meanings. We conclude with an optimistic outlook for logic-based provenance tools in next-generation biodiversity and phylogeny data platforms.
Motivation & Objective
- To address the challenge of tracking taxonomic changes across editions of global classification systems like MSW, where name usages may shift without clear provenance.
- To develop a method for representing and reasoning over taxonomic provenance that is both machine-processable and semantically transparent.
- To enable quantitative analysis of name/meaning dissociation in biodiversity data by aligning two influential primate classifications.
- To demonstrate that logic-based reasoning can produce a single, consistent alignment of taxonomic concepts across editions, even with significant changes.
Proposed method
- Individuating taxonomic name usages as discrete taxonomic concepts to enable precise comparison.
- Articulating relationships between concepts using human-asserted Region Connection Calculus (RCC-5) relations to model spatial and hierarchical overlaps.
- Employing an Answer Set Programming (ASP) toolkit to infer logically consistent alignments from the input articulations.
- Generating a comprehensive map of 153,111 Maximally Informative Relations to represent all concept pair provenance relationships.
- Validating the alignment by ensuring consistency across all input constraints and producing a single, coherent solution.
- Visualizing the alignment and its partitions to support further quantitative analysis of taxonomic congruence.
Experimental results
Research questions
- RQ1How can taxonomic provenance be systematically represented and reasoned over across multiple classification hierarchies?
- RQ2To what extent do name usages in MSW2 and MSW3 primates maintain semantic congruence across editions?
- RQ3Can logic-based inference resolve ambiguities in taxonomic changes between successive classification editions?
- RQ4What proportion of name usages exhibit name/meaning dissociation across the two MSW editions?
- RQ5Can automated, machine-processable tools improve the traceability and reliability of taxonomic data in biodiversity and phylogenetic platforms?
Key findings
- The reasoning process produced a single, logically consistent alignment of 483 MSW3 primate concepts against 317 MSW2 concepts using 402 input articulations.
- A total of 153,111 Maximally Informative Relations were inferred, forming a comprehensive provenance resolution map for all concept pairs.
- Approximately 33% of name usages across the two editions are not reliable in the sense that the same name does not identify congruent taxonomic meanings.
- The method successfully resolved complex taxonomic changes, including splits, lumpings, and reclassifications, with formal logical consistency.
- The approach demonstrates feasibility for next-generation biodiversity platforms to track taxonomic provenance at scale.
- The results highlight the need for granular, machine-processable identifiers to improve data interoperability and reproducibility in phylogenetic and biodiversity research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.