[Paper Review] Towards a semantic approach in GLAM Labs: the case of the Data Foundry at the National Library of Scotland
This paper presents a framework to transform GLAM (Galleries, Libraries, Archives, and Museums) metadata datasets into Linked Open Data (LOD) using Semantic Web standards, applying the method to three datasets from the National Library of Scotland's Data Foundry. The approach enables machine-readable, interoperable data publishing, with publicly available Jupyter Notebooks and results for reuse in digital humanities and data science.
GLAM organisations have been exploring the benefits of publishing their digital collections in a wide variety of forms since the 2000s. Many institutions, and in particular libraries, have adopted the Semantic Web and Linked Data principles to their main catalogues. Recent advances in technology and innovative approaches concerning the reuse of the digital collections by means of computational access have paved the way for the creation of Labs within GLAM organisations. In this work, we present a framework to transform the datasets made available by GLAM organisations under open licenses into LOD. The framework has been applied to three metadata datasets made available by the Data Foundry at the National Library of Scotland. The results of this work are publicly available and can be applied to other domains such as digital humanities and data science.
Motivation & Objective
- Address the challenge of making GLAM digital collections computationally accessible and interoperable through semantic publishing.
- Overcome limitations of traditional formats like MARC and CSV by converting them into machine-readable, standardized LOD.
- Provide a reusable, best-practice framework for GLAM institutions to publish their datasets as LOD.
- Demonstrate the feasibility and benefits of semantic enrichment using real-world datasets from the National Library of Scotland.
- Support international collaboration by aligning with the Linked Open Data Cloud and Wikidata ecosystems.
Proposed method
- Design a transformation pipeline that maps existing metadata (e.g., MARC, Dublin Core) to RDF using standardized vocabularies and URIs.
- Apply best practices from the W3C Linked Data guidelines and VoID vocabulary for dataset description and interlinking.
- Utilize Jupyter Notebooks to implement and document the transformation process, ensuring reproducibility.
- Employ tools like RDF/JS and XPath-based mappings to convert structured data (e.g., XML) into RDF triples.
- Integrate external knowledge sources (e.g., Wikidata, DBPedia) to enrich entities and improve data quality.
- Validate data quality through automated checks and alignment with established LOD principles.

Experimental results
Research questions
- RQ1How can traditional GLAM metadata formats be systematically transformed into machine-readable, semantically enriched Linked Open Data?
- RQ2What technical and methodological framework enables reliable, reproducible, and scalable LOD publishing from heterogeneous GLAM datasets?
- RQ3To what extent does semantic enrichment improve data interoperability and reusability in digital humanities and data science applications?
- RQ4What are the practical challenges and benefits of adopting LOD standards in institutional GLAM data publishing workflows?
- RQ5How can the resulting LOD datasets be effectively documented, published, and integrated into global knowledge graphs like the Linked Open Data Cloud?
Key findings
- The framework successfully transformed three metadata datasets from the National Library of Scotland’s Data Foundry into structured, machine-readable RDF format.
- The resulting LOD datasets are publicly available and interlinked with external knowledge bases such as Wikidata and DBPedia, enhancing their contextual richness.
- The use of Jupyter Notebooks enabled full reproducibility of the transformation process, supporting transparency and reuse by other institutions.
- The approach demonstrated significant improvements in data expressivity and machine processability compared to native MARC or CSV formats.
- The transformation pipeline adhered to W3C best practices for LOD publishing, including proper URI design and VoID metadata description.
- The study confirms that semantic publishing is feasible and beneficial for GLAM institutions seeking to increase the impact and reusability of their digital collections.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.