Skip to main content
QUICK REVIEW

[Paper Review] Automatic Identification of Research Fields in Scientific Papers

Éric Kergosien, Amin Farvardin|arXiv (Cornell University)|Jun 8, 2018
Geographic Information Systems Studies10 references3 citations
TL;DR

This paper presents a natural language processing and text mining framework for automatically identifying spatial, temporal, and thematic research fields in multilingual scientific papers. By standardizing heterogeneous corpora from ISTEX, CIRAD, and ANRT, the approach uses rule-based and machine learning techniques to annotate spatial entities with 90% precision and 60% recall, enabling geospatially aware scientific literature retrieval via a web-based tool (SISO).

ABSTRACT

The TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers. The project is divided into three main work packages: (1) identification of the periods and places of empirical studies, and which reflect the publications resulting from the analyzed text samples, (2) identification of the themes which appear in these documents, and (3) development of a web-based geographical information retrieval tool (GIR). The first two actions combine Natural Language Processing patterns with text mining methods. The integration of the spatial, thematic and temporal dimensions in a GIR contributes to a better understanding of what kind of research has been carried out, of its topics and its geographical and historical coverage. Another originality of the TERRE-ISTEX project is the heterogeneous character of the corpus, including PhD theses and scientific articles from the ISTEX digital libraries and the CIRAD research center.

Motivation & Objective

  • To develop a method for extracting and standardizing spatial, temporal, and thematic information from heterogeneous scientific literature.
  • To enable researchers to retrieve papers based on specific geographic territories, time periods, and research topics.
  • To support the analysis of disciplinary and multidisciplinary research trends in climate change studies across Senegal and Madagascar.
  • To build a scalable, multilingual, and extensible system for indexing and querying scientific corpora with geospatial and thematic metadata.
  • To create a web-based tool (SISO) that supports expert validation and annotation of spatial, temporal, and thematic entities in scientific documents.

Proposed method

  • Standardized multilingual scientific corpora from ISTEX, CIRAD (Agritrop), and ANRT (theses.fr) using a unified data model.
  • Applied rule-based and pattern-matching techniques in GATE (General Architecture for Text Engineering) to identify spatial, temporal, and thematic entities.
  • Used the Agrovoc ontology and offline lookup for thematic entity annotation, with custom lexicons for spatial and temporal features.
  • Developed the SISO web application to upload, visualize, correct, and export annotated corpora in XML/MODS format.
  • Integrated Lucene-based Elasticsearch for efficient indexing and geospatial information retrieval from the annotated corpus.
  • Evaluated performance on 140 manually annotated documents, achieving 78% precision in spatial entity annotation, with ongoing evaluation on 600 articles.

Experimental results

Research questions

  • RQ1How can spatial, temporal, and thematic information be automatically extracted from heterogeneous, multilingual scientific papers?
  • RQ2What is the performance of rule-based NLP techniques in identifying geographic locations and research topics in scientific literature?
  • RQ3Can a web-based annotation and validation system (SISO) effectively support domain experts in curating large-scale scientific corpora?
  • RQ4How does the integration of standardized metadata and semantic ontologies improve geospatial information retrieval in scientific literature?
  • RQ5To what extent can machine learning enhance disambiguation of spatial entities (e.g., distinguishing countries from organizations) in scientific texts?

Key findings

  • The system achieved 90% precision and 60% recall in spatial entity annotation on a 10-article English corpus, with 78% precision on a larger 140-document manual evaluation.
  • The entire processing pipeline for 8,500 documents took 16,105 seconds (about 1.9 seconds per document), demonstrating scalability.
  • The SISO web application enables efficient visualization, correction, and export of annotated corpora in XML/MODS format for downstream use.
  • The integration of the Agrovoc ontology and custom lexicons improved thematic entity annotation, particularly for agricultural and environmental topics.
  • The use of Elasticsearch for indexing enables efficient geospatial and thematic search over large corpora, supporting the development of a geographical information retrieval (GIR) system.
  • Future work will integrate machine learning (SVM-based models) to improve spatial entity disambiguation, with planned comparison to the ISO-Space model.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.