[论文解读] High-Performance Reachability Query Processing under Index Size Restrictions
本文提出 FERRARI,一种面向大规模有向图在严格索引大小限制下高性能可达性查询处理的自适应空间索引结构。通过选择性区间合并与引导式在线搜索实现传递闭包的自适应压缩,FERRARI 在真实世界规模的网页图上实现了亚微秒级查询响应时间,且在正向可达性查询上显著优于 GRAIL 等先前方法。
In this paper, we propose a scalable and highly efficient index structure for the reachability problem over graphs. We build on the well-known node interval labeling scheme where the set of vertices reachable from a particular node is compactly encoded as a collection of node identifier ranges. We impose an explicit bound on the size of the index and flexibly assign approximate reachability ranges to nodes of the graph such that the number of index probes to answer a query is minimized. The resulting tunable index structure generates a better range labeling if the space budget is increased, thus providing a direct control over the trade off between index size and the query processing performance. By using a fast recursive querying method in conjunction with our index structure, we show that in practice, reachability queries can be answered in the order of microseconds on an off-the-shelf computer - even for the case of massive-scale real world graphs. Our claims are supported by an extensive set of experimental results using a multitude of benchmark and real-world web-scale graph datasets.
研究动机与目标
- 解决在索引大小受限且主内存有限的超大规模图中高效可达性查询处理的挑战。
- 克服现有大小受限索引在正向可达性查询上的性能瓶颈,此类查询通常需要代价高昂的递归探索。
- 通过在用户定义的空间预算下灵活分配近似可达区间,实现索引大小与查询性能之间的可调制权衡。
- 设计一种索引结构,在保持高准确率与可扩展性的同时,最小化预期查询处理时间,适用于真实世界数据集。
提出的方法
- 该方法采用节点区间标记方案,将可达顶点编码为标识符范围,以紧凑方式表示可达集合。
- 在索引构建过程中,当超出空间预算时,自适应地合并相邻区间,以平衡压缩效率与查询效率。
- 将区间分配建模为区间覆盖问题,通过高效求解以在大小约束下最小化查询成本。
- 结合索引使用快速递归查询机制,实现引导式搜索,加速正向与负向可达性查询。
- 索引支持精确与近似可达区间,可通过增加空间预算灵活扩展性能。
- 构建阶段采用启发式策略,优先选择能减少每次查询预期索引探测次数的区间合并方案。
实验结果
研究问题
- RQ1能否构建一个空间受限的可达性索引,在保持高准确率的同时最小化预期查询处理时间?
- RQ2与现有方法相比,自适应区间压缩对正向可达性查询性能有何影响?
- RQ3该索引结构在多样化的真实世界图工作负载下,其查询速度与存储效率可扩展到何种程度?
- RQ4引导式在线搜索过程是否能显著减少可达性评估所需的索引探测次数?
主要发现
- 在 YAGO2 数据集上,FERRARI-L 对 10 万次随机查询的中位查询时间为 10.45ms,优于 GRAIL 的 30.53ms,查询时间减少 65.7%。
- 在 Twitter 数据集的正向查询中,FERRARI-L 将中位查询时间降低至 9.80ms,相比 GRAIL 的 18.21ms 提升 46.1%。
- 在 GovWild 数据集上,FERRARI-L 对正向查询的中位查询时间为 13.33ms,显著低于 GRAIL 的 29.84ms,减少 55.3%。
- FERRARI 保持了紧凑的索引大小,YAGO2 的索引大小为 5,844.87KB,远低于 GRAIL 的 59,587KB,同时性能更优。
- FERRARI 的构建时间在所有数据集上均显著低于 GRAIL,其中 FERRARI-L 在 YAGO2 上仅耗时 137.88ms,而 GRAIL 为 182.96ms。
- 在 Web-UK 数据集上,FERRARI-L 对 10 万次正向查询的中位查询时间为 16.85ms,相比 GRAIL 的 26,792.72ms 提升超过 1,500 倍。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。