Skip to main content
QUICK REVIEW

[Paper Review] A Reproducible Journal Classification and Global Map of Science Based on Aggregated Journal-Journal Citation Relations.

Loet Leydesdorff, Lutz Bornmann|arXiv (Cornell University)|Apr 10, 2016
Advanced Text Analysis Techniques3 citations
TL;DR

This paper proposes a reproducible, non-subjective journal classification system based on aggregated journal-journal citation data from 11,149 journals in the Science and Social Science Citation Indexes. Using VOSviewer and Pajek to generate a hierarchical, tree-like classification, it produces a stable, reproducible map of science with nine top-level fields, including a detailed decomposition of the social sciences into LIS and STS, enabling objective comparison of scientific fields over time.

ABSTRACT

A number of journal classification systems have been developed in bibliometrics since the launch of the Citation Indices by the Institute of Scientific Information (ISI) in the 1960s. The best known system is the so-called Web-of-Science Subject Categories (WCs). Each system has its own advantages and disadvantages. Using the Journal Citation Reports 2014 of the Science Citation Index and the Social Science Citation Index (n of journals = 11,149), we examine the options for developing an unambiguous classification of the journals into subject categories on the basis of aggregated journal-journal citation data. Combining routines in VOSviewer and Pajek, a tree-like classification is developed which can be reproduced unambiguously. At each level one can generate a map of science for all the journals subsumed under the category. Nine major fields are distinguished at the top level, with the social sciences as a single field (n of journals = 3,131). In this study, further decomposition of the social sciences is pursued for the sake of example with a focus on journals in information science (LIS) and science studies (STS). The classification improves alternative options by removing subjective judgement and avoiding the problem of randomness in the seed number that has made algorithmic solutions hitherto irreproducible. A non-subjective map and classification can provide a baseline for measuring the effectiveness of policy interventions. Maps can be compared between years, and change in the (social) sciences can be indicated.

Motivation & Objective

  • To develop a reproducible, non-subjective journal classification system based on citation data to overcome the limitations of existing, often arbitrary or random seed-based methods.
  • To create a stable, hierarchical classification of journals into subject categories using aggregated journal-journal citation relations.
  • To generate a global map of science that can be reproduced and compared across time, providing a baseline for measuring policy effectiveness.
  • To demonstrate the feasibility of decomposing the social sciences into subfields like LIS and STSTS using citation-based clustering.
  • To eliminate subjective judgment and randomness in algorithmic classification by using deterministic, data-driven routines.

Proposed method

  • Aggregating journal-journal citation data from the Journal Citation Reports 2014 of the Science Citation Index and Social Science Citation Index (n = 11,149 journals).
  • Applying network analysis techniques in Pajek to identify dense citation clusters among journals.
  • Using VOSviewer to visualize and refine the hierarchical structure of citation clusters into a tree-like classification.
  • Implementing deterministic clustering routines that avoid random seed selection, ensuring reproducibility.
  • Building a multi-level classification with nine top-level fields, including social sciences as one field, and further decomposing it for example subfields.
  • Generating science maps at each classification level to visualize the structure of scientific domains.

Experimental results

Research questions

  • RQ1Can a reproducible, non-subjective journal classification be developed using aggregated journal-journal citation data?
  • RQ2How can citation-based clustering be structured to produce a stable, hierarchical classification of journals without subjective input?
  • RQ3To what extent can the social sciences be meaningfully decomposed into subfields like LIS and STS using citation data?
  • RQ4Can this classification system serve as a baseline for measuring changes in scientific fields over time?
  • RQ5How does this method improve upon existing algorithmic classification systems that suffer from randomness due to seed dependency?

Key findings

  • The proposed method produces a reproducible, deterministic classification of 11,149 journals into nine major fields, with the social sciences forming one of these fields (n = 3,131 journals).
  • The classification process successfully decomposes the social sciences into subfields such as LIS and STS, demonstrating the method's granularity and applicability.
  • By avoiding random seed selection, the method eliminates the irreproducibility issue common in prior algorithmic classification approaches.
  • The resulting classification enables the generation of stable, repeatable maps of science at each hierarchical level, facilitating longitudinal comparison.
  • The approach provides a non-subjective baseline for assessing changes in scientific fields and evaluating policy impacts over time.
  • The integration of VOSviewer and Pajek allows for both robust clustering and clear visualization of the resulting science maps.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.