Skip to main content
QUICK REVIEW

[论文解读] A Stitch in Time Saves Nine -- SPARQL querying of Property Graphs using Gremlin Traversals

Harsh Thakkar, Dharmen Punjani|arXiv (Cornell University)|Jan 9, 2018
Graph Theory and Algorithms参考文献 34被引用 13
一句话总结

本文提出 Gremlinator,一种新颖的 SPARQL 到 Gremlin 翻译器,使 SPARQL 查询能够通过 Gremlin 遍历在属性图数据库上执行。通过将 SPARQL 基本图模式(BGP)翻译为声明式 Gremlin 模式匹配遍历,该方法实现了显著的性能提升——尤其在星型结构和复杂查询中表现突出,相较于 Virtuoso 和 4Store 等 RDF 存储中的原生 SPARQL 引擎,展现出显著的速度优势。

ABSTRACT

Knowledge graphs have become popular over the past years and frequently rely on the Resource Description Framework (RDF) or Property Graphs (PG) as underlying data models. However, the query languages for these two data models -- SPARQL for RDF and Gremlin for property graph traversal -- are lacking interoperability. We present Gremlinator, a novel SPARQL to Gremlin translator. Gremlinator translates SPARQL queries to Gremlin traversals for executing graph pattern matching queries over graph databases. This allows to access and query a wide variety of Graph Data Management Systems (DMS) using the W3C standardized SPARQL query language and avoid the learning curve of a new Graph Query Language. Gremlin is a system-agnostic traversal language covering both OLTP graph database or OLAP graph processors, thus making it a desirable choice for supporting interoperability wrt. querying Graph DMSs. We present a comprehensive empirical evaluation of Gremlinator and demonstrate its validity and applicability by executing SPARQL queries on top of the leading graph stores Neo4J, Sparksee, and Apache TinkerGraph and compare the performance with the RDF stores Virtuoso, 4Store and JenaTDB. Our evaluation demonstrates the substantial performance gain obtained by the Gremlin counterparts of the SPARQL queries, especially for star-shaped and complex queries.

研究动机与目标

  • 解决 SPARQL(RDF)与 Gremlin(属性图)查询语言之间的互操作性不足问题。
  • 使 SPARQL 用户无需学习新语言即可查询属性图数据库。
  • 通过利用属性图数据库的局部性感知遍历能力,提升查询性能。
  • 为不同图数据管理系统的统一查询接口提供支持。
  • 证明将 SPARQL 翻译为 Gremlin 以实现混合图查询执行的可行性与性能优势。

提出的方法

  • 通过将 SPARQL 操作逐步映射为 Gremlin 遍历步骤,将 SPARQL 基本图模式(BGPs)映射为声明式 Gremlin 模式匹配遍历。
  • 以 Apache TinkerPop 框架作为目标执行环境,确保与多种属性图数据库(如 Neo4j、Sparksee、TinkerGraph)的兼容性。
  • 将每个 SPARQL 查询翻译为相应的 Gremlin 遍历,保持语义等价性的同时优化图局部性。
  • 采用与系统无关的方法,避免使用特定供应商的 Gremlin 语法,确保在 OLTP 和 OLAP 图处理器之间的可移植性。
  • 设计翻译器以支持命令式与声明式遍历风格,重点聚焦于声明式模式匹配,以实现公平的性能比较。
  • 通过在 BSBM 和 Northwind 数据集上对多种 DMS(包括 RDF 存储(Virtuoso、4Store、JenaTDB)和属性图存储(Neo4j、Sparksee、TinkerGraph))进行实证评估,验证翻译正确性。

实验结果

研究问题

  • RQ1SPARQL 查询能否被有效翻译为 Gremlin 遍历,以实现 RDF 与属性图系统之间的互操作性?
  • RQ2将 SPARQL 翻译为 Gremlin 是否能在 RDF 存储上的原生 SPARQL 执行中带来可测量的性能提升?
  • RQ3对于星型结构和复杂图模式,SPARQL 与 Gremlin 的查询性能特征有何差异?
  • RQ4Gremlinator 是否能实现在异构图数据库上的无缝查询执行,而无需修改应用程序?
  • RQ5对于不同类型的查询,使用 Gremlin 遍历相较于原生 SPARQL 引擎的性能开销或优势如何?

主要发现

  • Gremlinator 能够成功将 SPARQL 查询翻译为等价的 Gremlin 模式匹配遍历,实现 RDF 与属性图系统之间的互操作性。
  • 对于星型结构和复杂查询,Gremlin 遍历的性能相比原生 SPARQL 执行平均提升最高达 3.5 倍,部分查询甚至超过 10 倍。
  • 在 BSBM 数据集上,Gremlin 遍历在 TinkerGraph 上的执行时间为 11.49–17.01 ms(无索引),而 Virtuoso 上为 36.25–52 ms(有索引),性能优势显著。
  • 在 Northwind 数据集上,Gremlinator 在 TinkerGraph 上实现最低 0.39 ms(L2 查询)的执行时间,而 4Store 上为 113 ms,凸显了对特定模式的卓越性能。
  • 性能优势在受益于图局部性的查询中最为明显,此时 Gremlin 的微索引和遍历效率优于 SPARQL 的基于连接的执行方式。
  • 评估结果证实,Gremlinator 能够在包括 OLTP 和 OLAP 系统在内的多种图数据库上实现高效、可移植且高性能的 SPARQL 查询。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。