[Paper Review] Criteria-first, semantics-later: reproducible structure discovery in image-based sciences
The paper proposes a criteria-first, semantics-later framework that extracts semantics-free structural products from measurements under explicit criteria, enabling downstream, plural semantic mappings and improved reproducibility across image-based sciences.
Across the natural and life sciences, images have become a primary measurement modality, yet the dominant analytic paradigm remains semantics-first. Structure is recovered by predicting or enforcing domain-specific labels. This paradigm fails systematically under the conditions that make image-based science most valuable, including open-ended scientific discovery, cross-sensor and cross-site comparability, and long-term monitoring in which domain ontologies and associated label sets drift culturally, institutionally, and ecologically. A deductive inversion is proposed in the form of criteria-first and semantics-later. A unified framework for criteria-first structure discovery is introduced. It separates criterion-defined, semantics-free structure extraction from downstream semantic mapping into domain ontologies or vocabularies and provides a domain-general scaffold for reproducible analysis across image-based sciences. Reproducible science requires that the first analytic layer perform criterion-driven, semantics-free structure discovery, yielding stable partitions, structural fields, or hierarchies defined by explicit optimality criteria rather than local domain ontologies. Semantics is not discarded; it is relocated downstream as an explicit mapping from the discovered structural product to a domain ontology or vocabulary, enabling plural interpretations and explicit crosswalks without rewriting upstream extraction. Grounded in cybernetics, observation-as-distinction, and information theory's separation of information from meaning, the argument is supported by cross-domain evidence showing that criteria-first components recur whenever labels do not scale. Finally, consequences are outlined for validation beyond class accuracy and for treating structural products as FAIR, AI-ready digital objects for long-term monitoring and digital twins.
Motivation & Objective
- Identify limitations of semantics-first pipelines in open-ended discovery, cross-site comparability, and long-term monitoring.
- Propose a domain-general, criteria-first framework that yields stable, semantics-free structural products.
- Demonstrate how downstream semantics can map the same structure to multiple ontologies without rewriting upstream extraction.
- Argue for treating structural products as FAIR, AI-ready digital objects suitable for digital twins and long-term monitoring.
Proposed method
- Formalize a minimal framework: a measurement field X, an explicit criterion C, and a structure-extraction operator S_C yielding a structural product S.
- Define S_C as the solution to an optimization problem maximizing a criterion-derived objective E_C(X,S) over admissible structures S.
- Allow multiple structure types (partitions, graphs, hierarchies, structural fields) depending on modality and criterion.
- Introduce downstream semantic mappings M_i: S -> O_i to map the structural product to domain ontologies, enabling plural interpretations.
- Provide postulates for reproducibility: explicit criterion, determinacy, stability under perturbations, and mapping pluralism.
- Illustrate the approach with cross-domain evidence showing the recurring emergence of criterion-first components when labels do not scale.
Experimental results
Research questions
- RQ1What are the limitations of semantics-first pipelines for long-term monitoring and open-ended discovery?
- RQ2How can a criterion-defined, semantics-free structural layer be designed to be reproducible and transferable across domains?
- RQ3How should downstream semantics be organized to map the same structural product to multiple domain ontologies without upstream rewrites?
- RQ4What evidence supports the ubiquity of the criteria-first pattern across image-based sciences?
Key findings
- A semantics-first approach is brittle under domain shift, cross-sensor variation, and evolving ontologies.
- A criterion-first layer yields stable, transferable structural products independent of domain labels.
- Downstream semantics can be plural and purpose-bound, mapping the same structure to different ontologies without altering the upstream extraction.
- Structural products can serve as durable, FAIR, AI-ready digital objects for digital twins and long-term monitoring.
- Evaluation should focus on robustness, scale coherence, and global consistency rather than fixed semantic accuracy.
- Self-supervised and foundation-model approaches can implement the criterion-first structure extraction while keeping explicit criteria testable and auditable.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.