[Paper Review] Informational Space of Meaning for Scientific Texts
This paper introduces a novel vector space model called the 'Meaning Space' that quantifies word meaning in scientific texts using Relative Information Gain (RIG) across 252 Web of Science subject categories. Applied to the Leicester Scientific Corpus (1.67M abstracts) and dictionary (LScDC), the RIG-based representation outperforms raw frequency in identifying topic-specific, high-impact scientific terms, with a 103,998 × 252 RIG matrix and a new scientific thesaurus (LScT) publicly available.
In Natural Language Processing, automatic extracting the meaning of texts constitutes an important problem. Our focus is the computational analysis of meaning of short scientific texts (abstracts or brief reports). In this paper, a vector space model is developed for quantifying the meaning of words and texts. We introduce the Meaning Space, in which the meaning of a word is represented by a vector of Relative Information Gain (RIG) about the subject categories that the text belongs to, which can be obtained from observing the word in the text. This new approach is applied to construct the Meaning Space based on Leicester Scientific Corpus (LSC) and Leicester Scientific Dictionary-Core (LScDC). The LSC is a scientific corpus of 1,673,350 abstracts and the LScDC is a scientific dictionary which words are extracted from the LSC. Each text in the LSC belongs to at least one of 252 subject categories of Web of Science (WoS). These categories are used in construction of vectors of information gains. The Meaning Space is described and statistically analysed for the LSC with the LScDC. The usefulness of the proposed representation model is evaluated through top-ranked words in each category. The most informative n words are ordered. We demonstrated that RIG-based word ranking is much more useful than ranking based on raw word frequency in determining the science-specific meaning and importance of a word. The proposed model based on RIG is shown to have ability to stand out topic-specific words in categories. The most informative words are presented for 252 categories. The new scientific dictionary and the 103,998 x 252 Word-Category RIG Matrix are available online. Analysis of the Meaning Space provides us with a tool to further explore quantifying the meaning of a text using more complex and context-dependent meaning models that use co-occurrence of words and their combinations.
Motivation & Objective
- To develop a computational model for quantifying word meaning in short scientific texts such as abstracts.
- To address the challenge of identifying topic-specific, semantically important words in scientific literature beyond simple frequency counts.
- To construct a Meaning Space where word meaning is represented by vectors of Relative Information Gain (RIG) across subject categories.
- To create a publicly available scientific dictionary and RIG matrix for use in NLP and text mining applications.
Proposed method
- Represent each word as a vector of Relative Information Gain (RIG) values across 252 Web of Science subject categories.
- Compute RIG for each word-category pair using information-theoretic principles to measure how much a word's presence reduces uncertainty about the category.
- Construct the Meaning Space using the Leicester Scientific Corpus (LSC) of 1,673,350 scientific abstracts and the Leicester Scientific Dictionary-Core (LScDC) derived from it.
- Use the RIG matrix to rank the most informative words per category, enabling identification of domain-specific terminology.
- Apply dimensionality reduction and clustering techniques to explore the structure of the Meaning Space.
- Release the 103,998 × 252 Word-Category RIG Matrix and a new scientific thesaurus (LScT) for public use.
Experimental results
Research questions
- RQ1Can RIG-based word representation outperform raw frequency in identifying the most meaningful and topic-specific words in scientific texts?
- RQ2How does the RIG-based vector space model capture the semantic specificity of scientific terminology across diverse research fields?
- RQ3To what extent does the Meaning Space constructed via RIG reveal meaningful patterns in word usage and category associations?
- RQ4Can the RIG matrix and resulting thesaurus (LScT) serve as a robust, data-driven tool for scientific text analysis and information extraction?
Key findings
- RIG-based ranking significantly outperforms raw frequency in identifying topic-specific, high-impact scientific terms across all 252 subject categories.
- The most informative words in each category—such as 'femal' (RIG: 3.6×10⁻²) in Women’s Studies and 'speci' (RIG: 1.9×10⁻¹) in Zoology—were highly relevant and contextually meaningful.
- The 103,998 × 252 Word-Category RIG Matrix provides a comprehensive, publicly accessible representation of scientific word meaning across disciplines.
- The proposed Meaning Space enables effective discovery of anomalies and domain-specific terminology through information gain analysis.
- The Leicester Scientific Thesaurus (LScT) was successfully generated from the RIG matrix, offering a new resource for scientific text mining.
- The model demonstrates strong potential for use in advanced meaning models involving word co-occurrence and semantic combinations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.