Skip to main content
QUICK REVIEW

[Paper Review] Data Mesh: a Systematic Gray Literature Review

Abel Goedegebuure, Indika Kumara|arXiv (Cornell University)|Apr 3, 2023
Data Quality and Management7 citations
TL;DR

This paper conducts a systematic gray literature review of 114 sources to define data mesh, identify its four principles, map findings to SOA reference architectures, and outline research challenges.

ABSTRACT

Data mesh is an emerging domain-driven decentralized data architecture that aims to minimize or avoid operational bottlenecks associated with centralized, monolithic data architectures in enterprises. The topic has picked the practitioners' interest, and there is considerable gray literature on it. At the same time, we observe a lack of academic attempts at defining and building upon the concept. Hence, in this article, we aim to start from the foundations and characterize the data mesh architecture regarding its design principles, architectural components, capabilities, and organizational roles. We systematically collected, analyzed, and synthesized 114 industrial gray literature articles. The review provides insights into practitioners' perspectives on the four key principles of data mesh: data as a product, domain ownership of data, self-serve data platform, and federated computational governance. Moreover, due to the comparability of data mesh and SOA (service-oriented architecture), we mapped the findings from the gray literature into the reference architectures from the SOA academic literature to create the reference architectures for describing three key dimensions of data mesh: organization of capabilities and roles, development, and runtime. Finally, we discuss open research issues in data mesh, partially based on the findings from the gray literature.

Motivation & Objective

  • Define data mesh and its four guiding principles from practitioner literature (data as a product, domain ownership of data, self-serve data platform, federated computational governance).
  • Identify benefits, concerns, and organizational applicability of data mesh in industry settings.
  • Develop reference architectures for data mesh by mapping gray-literature findings to SOA architectures (organization of capabilities/roles, development, runtime).
  • Highlight open research challenges in data mesh to guide academic investigation.

Proposed method

  • Followed gray literature review guidelines adapted for GLRs (Garousi et al., and Kitchenham/Charters) to ensure systematicity.
  • Collected 114 gray literature sources (2019–2022) using Google search with defined queries: "Data Mesh" and "Decentralized Data Architecture".
  • Applied inclusion/exclusion and quality criteria; performed inter-rater reliability with Cohen’s Kappa = 0.79.
  • Used structural and descriptive coding in Atlas.ti for qualitative analysis; iteratively developed categories and themes.
  • Mapped gray-literature findings to service-oriented architecture (SOA) reference architectures to construct three data mesh reference architectures.
  • Identified open research issues by linking practitioner challenges to SOA and data management literature.
Figure 1. Google trends for the word “Data Mesh”.
Figure 1. Google trends for the word “Data Mesh”.

Experimental results

Research questions

  • RQ1RQ1: What is Data Mesh and why is it needed, focusing on the four design principles (data as a product, domain ownership, self-serve platform, federated governance).
  • RQ2RQ2: What are the benefits and concerns of adopting data mesh and its implementation challenges?
  • RQ3RQ3: How should an organization build data mesh, and can reference architectures be established by mapping to SOA architectures?
  • RQ4RQ4: What are the research challenges concerning data mesh based on practitioner literature?

Key findings

  • Data mesh is a domain-oriented decentralized architecture for managing analytical data at scale, decomposing a monolithic data space into domain-aligned data products.
  • Data products have components (data, metadata, code, interfaces, infrastructure) and two types (atomic and composite).
  • Eight characteristics define high-quality data products (discoverable, interoperable, natively accessible, self-describing, understandable, secure, trustworthy, valuable).
  • Federated computational governance balances global standards with local domain autonomy to ensure interoperability and compliance.
  • A self-serve data platform and platform tooling support domain teams in building and operating data products; automation and governance are central to scalability.
  • The gray literature is mapped to three SOA-based reference architectures, aiding organizational design in development and runtime aspects, and highlighting research challenges.
Figure 2. Systematic gray literature review process.
Figure 2. Systematic gray literature review process.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.