Skip to main content
QUICK REVIEW

[Paper Review] Fostering the integration of European Open Data into Data Spaces through High-Quality Metadata

Javier Conde, Alejandro Pozo|arXiv (Cornell University)|Feb 8, 2024
Data Quality and ManagementDecision Sciences3 citations
TL;DR

This paper presents an open-source software stack that automates the generation, validation, and publication of high-quality, DCAT-compliant metadata for European Open Data, enabling seamless integration into Data Spaces. The solution, validated via the YODA portal with over 200 datasets, achieved a perfect 100/100 score for findability and top-tier performance across FAIR metrics on data.europa.eu, demonstrating scalability and interoperability in real-world deployment.

ABSTRACT

The term Data Space, understood as the secure exchange of data in distributed systems, ensuring openness, transparency, decentralization, sovereignty, and interoperability of information, has gained importance during the last years. However, Data Spaces are in an initial phase of definition, and new research is necessary to address their requirements. The Open Data ecosystem can be understood as one of the precursors of Data Spaces as it provides mechanisms to ensure the interoperability of information through resource discovery, information exchange, and aggregation via metadata. However, Data Spaces require more advanced capabilities including the automatic and scalable generation and publication of high-quality metadata. In this work, we present a set of software tools that facilitate the automatic generation and publication of metadata, the modeling of datasets through standards, and the assessment of the quality of the generated metadata. We validate all these tools through the YODA Open Data Portal showing how they can be connected to integrate Open Data into Data Spaces.

Motivation & Objective

  • To address the critical barrier of low-quality or non-compliant metadata in European Open Data portals, which hinders integration into emerging Data Spaces.
  • To enable automated, scalable, and standardized metadata generation aligned with DCAT-AP v2.1.0 and FAIR principles.
  • To develop a reusable software stack that transforms raw datasets into publishable, semantically rich metadata for interoperable data exchange.
  • To validate the solution in real-world scenarios across diverse public data sources (e.g., AEMET, SmartSantander, Smart Campus CEI Moncloa).
  • To demonstrate that high-quality metadata significantly improves data findability, accessibility, and reusability in pan-European data infrastructures.

Proposed method

  • Leveraged Apache NiFi for scalable, real-time data ingestion and transformation from heterogeneous data sources.
  • Implemented a metadata generation pipeline that maps raw datasets to the DCAT-AP v2.1.0 standard using semantic modeling and metadata templates.
  • Integrated a metadata quality assessment module based on the Open Data Portal Quality Assessment (MQA) framework to evaluate findability, accessibility, interoperability, reusability, and contextuality.
  • Automated the publication of datasets to CKAN-based portals using the ckanext-dcatapedp extension with OAI-PMH v2 harvesting support.
  • Designed a modular, open-source software stack enabling reuse across different public sector data providers and integration with the European Data Portal.
  • Configured the system for three distinct data sources (AEMET, SmartSantander, Smart Campus CEI Moncloa) to validate scalability and cross-domain applicability.

Experimental results

Research questions

  • RQ1How can automated, high-quality metadata generation be achieved at scale for heterogeneous public sector datasets?
  • RQ2To what extent does DCAT-AP v2.1.0 compliance improve the FAIRness and interoperability of Open Data in Data Spaces?
  • RQ3Can a standardized, automated pipeline significantly enhance metadata quality metrics such as findability and reusability in real-world deployments?
  • RQ4What is the impact of automated metadata quality assessment on the performance of Open Data portals in pan-European data infrastructures?
  • RQ5How effectively can the proposed stack integrate diverse public data sources into the European Data Portal while maintaining high metadata quality?

Key findings

  • The YODA portal achieved a perfect 100.0/100.0 score for findability on data.europa.eu, significantly outperforming the average of 69.3 among other harvested catalogs.
  • The portal scored 96.0/100 for accessibility, compared to an average of 53.7, demonstrating robust data availability and machine-readability.
  • Interoperability was rated at 80.0/100, far exceeding the average of 40.2, indicating strong DCAT-AP compliance and structured metadata.
  • Reusability scored 75.0/100, surpassing the average of 51.3, reflecting effective metadata richness and data documentation.
  • Contextuality, a key dimension for data discovery, scored 20.0/100—more than three times the average of 6.6—indicating superior data context and provenance information.
  • The solution successfully ingested and published over 200 datasets across three distinct public data sources (AEMET, SmartSantander, Smart Campus CEI Moncloa), proving scalability and cross-domain compatibility.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.