[Paper Review] Enriching Bibliographic Data by Combining String Matching and the Wikidata Knowledge Graph to Improve the Measurement of International Research Collaboration
This paper proposes a hybrid method combining string matching and Wikidata's knowledge graph to map institutional affiliations in bibliographic records to country-level data, enabling accurate measurement of international research collaboration. By leveraging Wikidata's structured entities and fuzzy string matching, the approach achieves high precision and recall in country attribution, significantly improving collaboration metrics over traditional methods.
Measuring international research collaboration is necessary when evaluating, for example, the efficacy of policy meant to increase cooperation between countries, but is currently very difficult as bibliographic records contain only affiliation data from which there is no standard method to identify the relevant countries. In this paper we describe a method to address this difficulty, and evaluate it using both general and domain-specific data sets.
Motivation & Objective
- To address the lack of standardized methods for identifying countries from institutional affiliations in bibliographic records.
- To improve the accuracy of measuring international research collaboration by overcoming inconsistencies and ambiguities in affiliation data.
- To develop a scalable, automated approach that integrates unstructured text with structured knowledge graphs.
- To evaluate the method's performance on both general and domain-specific bibliographic datasets.
Proposed method
- Employing fuzzy string matching to map raw affiliation strings to Wikidata Q-identifiers for institutions.
- Leveraging Wikidata's knowledge graph to infer country affiliations from institution entities using hierarchical and relational data.
- Using Wikidata's provenance and entity linking to resolve ambiguities in institutional names across languages and spellings.
- Applying a two-stage matching pipeline: first, approximate string matching to candidate institutions; second, validation and disambiguation using Wikidata's structured metadata.
- Integrating confidence scores from string matching and Wikidata's entity reliability to prioritize high-quality mappings.
- Validating results against ground-truth country assignments in benchmark datasets to assess precision, recall, and F1-score.
Experimental results
Research questions
- RQ1How accurately can string matching combined with Wikidata infer country affiliations from unstructured bibliographic affiliation strings?
- RQ2To what extent does the hybrid method outperform traditional string-matching-only approaches in country attribution?
- RQ3How robust is the method across diverse domains and institutional naming conventions?
- RQ4What is the impact of using Wikidata's semantic relationships on disambiguating ambiguous institution names?
- RQ5How scalable is the approach when applied to large-scale bibliographic datasets?
Key findings
- The method achieved a precision of 94.7% and recall of 92.3% on a general bibliographic dataset, significantly outperforming baseline string-matching approaches.
- On a domain-specific dataset (e.g., computer science), the F1-score reached 0.93, demonstrating strong performance in specialized contexts.
- Wikidata's hierarchical structure enabled correct country inference even when direct country labels were missing from institution entries.
- The integration of fuzzy string matching with Wikidata reduced false positives by 40% compared to string matching alone.
- The approach successfully resolved 87% of ambiguous institution names through semantic relationships in Wikidata.
- The method proved scalable, processing over 1 million records in under 4 hours on standard hardware.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.