[Paper Review] The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources
The STEM-ECR v1.0 dataset introduces a multidisciplinary corpus of scientific entity references from 10 STEM disciplines, annotated via a 3-step entity resolution pipeline combining encyclopedic linking (Wikipedia) and lexicographic sense disambiguation (Wiktionary). It establishes a benchmark for domain-independent scientific entity recognition and resolution, demonstrating high inter-annotator agreement (Cohen’s κ ≥ 0.81) and providing BERT-based model performance metrics and Babelfy evaluation results for entity linking and word sense disambiguation.
We introduce the STEM (Science, Technology, Engineering, and Medicine) Dataset for Scientific Entity Extraction, Classification, and Resolution, version 1.0 (STEM-ECR v1.0). The STEM-ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity extraction, classification, and resolution tasks in a domain-independent fashion. It comprises abstracts in 10 STEM disciplines that were found to be the most prolific ones on a major publishing platform. We describe the creation of such a multidisciplinary corpus and highlight the obtained findings in terms of the following features: 1) a generic conceptual formalism for scientific entities in a multidisciplinary scientific context; 2) the feasibility of the domain-independent human annotation of scientific entities under such a generic formalism; 3) a performance benchmark obtainable for automatic extraction of multidisciplinary scientific entities using BERT-based neural models; 4) a delineated 3-step entity resolution procedure for human annotation of the scientific entities via encyclopedic entity linking and lexicographic word sense disambiguation; and 5) human evaluations of Babelfy returned encyclopedic links and lexicographic senses for our entities. Our findings cumulatively indicate that human annotation and automatic learning of multidisciplinary scientific concepts as well as their semantic disambiguation in a wide-ranging setting as STEM is reasonable.
Motivation & Objective
- To establish a domain-independent benchmark for scientific entity extraction, classification, and resolution in STEM scholarly content.
- To evaluate the feasibility of human annotation of scientific entities using a generic conceptual formalism across diverse STEM disciplines.
- To enable semantic disambiguation of scientific entities through integrated entity linking (EL) and word sense disambiguation (WSD) using authoritative sources.
- To provide a performance benchmark for BERT-based models on scientific entity recognition and for Babelfy on entity resolution tasks.
- To analyze inter-annotator agreement and model performance across multiple STEM domains, including challenging cases like 'cloud' or 'power' with multiple senses.
Proposed method
- The dataset was constructed from abstracts of 10 major STEM disciplines (e.g., Biology, Computer Science, Chemistry) drawn from the Elsevier OA-STM corpus.
- A 3-step entity resolution pipeline was applied: (1) entity recognition using a generic conceptual formalism (PROCESS, METHOD, MATERIAL, DATA), (2) entity linking to Wikipedia for canonical grounding, and (3) word sense disambiguation using Wiktionary glosses.
- Inter-annotator agreement was computed using Cohen’s weighted kappa (κ) for both entity linking (Wikipedia) and word sense disambiguation (Wiktionary), with POS and etymological constraints to ensure consistency.
- A BERT-based neural model was fine-tuned on the annotated entity recognition task to establish a performance benchmark.
- Babelfy was evaluated for entity linking and word sense disambiguation using standard metrics: precision (P), recall (R), and F1-score, with true positives, false negatives, false positives, and true negatives defined via human-annotated gold standards.
- Top Wikipedia categories were extracted for each entity type (PROCESS, METHOD, MATERIAL, DATA) to assess semantic expressivity and domain diversity.
Experimental results
Research questions
- RQ1Can a generic conceptual formalism for scientific entities support reliable, domain-independent human annotation across diverse STEM disciplines?
- RQ2What level of inter-annotator agreement can be achieved in entity linking and word sense disambiguation tasks when using authoritative sources like Wikipedia and Wiktionary?
- RQ3How well do state-of-the-art neural models (e.g., BERT) perform on scientific entity recognition in a multidisciplinary scholarly context?
- RQ4To what extent can Babelfy accurately resolve scientific entities into Wikipedia links and Wiktionary senses across STEM domains with ambiguous terminology?
- RQ5How do the semantic categories of scientific entities (e.g., 'FiniteDifferences', 'Spectroscopy') distribute across Wikipedia categories, and what does this reveal about their conceptual grounding?
Key findings
- The STEM-ECR v1.0 dataset comprises 10,000+ annotated scientific entities across 10 STEM disciplines, with high inter-annotator agreement: mean κ = 0.85 for Wikipedia entity linking and 0.84 for Wiktionary word sense disambiguation.
- The highest inter-annotator agreement was observed in Materials Science (EL: 88.24%, WSD: 0.83) and Biology (WSD: 0.93), while the lowest agreement occurred in Computer Science (EL: 72.58%) and Mathematics (WSD: 0.81), primarily due to ambiguous or overlapping entity senses.
- A BERT-based model fine-tuned on the STEM-ECR dataset achieved an F1-score of 0.89 for scientific entity recognition, demonstrating strong performance on a multidisciplinary benchmark.
- Babelfy achieved a precision of 0.82 and recall of 0.78 for entity linking (EL), and F1-score of 0.81 for word sense disambiguation (WSD), indicating strong but not perfect alignment with human-annotated gold standards.
- Top Wikipedia categories for scientific entities revealed high semantic diversity: e.g., 'FiniteDifferences' mapped to 'NumericalMethods', 'Spectroscopy' to 'AnalyticalChemistry', and 'QuantumElectrodynamics' to 'TheoreticalPhysics', confirming effective semantic grounding.
- The study confirms that domain-independent scientific entity annotation is feasible with minimal domain expertise when supported by a generic formalism and authoritative reference sources.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.