Skip to main content
QUICK REVIEW

[论文解读] Machine Knowledge: Creation and Curation of Comprehensive Knowledge Bases

Gerhard Weikum, Xin Dong|arXiv (Cornell University)|Sep 24, 2020
Data Quality and Management参考文献 604被引用 115
一句话总结

对自动构建和整理大型知识库(KB)的方法的全面综述,涵盖实体发现、规范化、属性与关系提取、开放模式以及长期KB维护,并对主要KB进行案例研究。

ABSTRACT

Equipping machines with comprehensive knowledge of the world's entities and their relationships has been a long-standing goal of AI. Over the last decade, large-scale knowledge bases, also known as knowledge graphs, have been automatically constructed from web contents and text sources, and have become a key asset for search engines. This machine knowledge can be harnessed to semantically interpret textual phrases in news, social media and web tables, and contributes to question answering, natural language processing and data analytics. This article surveys fundamental concepts and practical methods for creating and curating large knowledge bases. It covers models and methods for discovering and canonicalizing entities and their semantic types and organizing them into clean taxonomies. On top of this, the article discusses the automatic extraction of entity-centric properties. To support the long-term life-cycle and the quality assurance of machine knowledge, the article presents methods for constructing open schemas and for knowledge curation. Case studies on academic projects and industrial knowledge graphs complement the survey of concepts and methods.

研究动机与目标

  • 让机器具备对世界知识的全面掌握,以支持 AI 应用的目标。
  • 对实体为中心的知识库及其生命周期的基础概念与体系结构设计进行综述。
  • 提出知识库创建的核心任务:发现、规范化、增强、开放模式演化以及整理。
  • 强调使用半结构化和非结构化来源的实际方法与设计决策。
  • 提供知名知识库项目的案例研究,以阐明原则与挑战。

提出的方法

  • 描述实体、类、属性及高阶关系的知识表示基础。
  • 概述知识库构建的设计空间,包括输入源、质量和输出范围。
  • 详细说明从多元来源进行实体发现和分类法构建的方法。
  • 解释实体规范化,包括实体连接与匹配。
  • 展示从文本和半结构化数据中提取属性与关系的技术。
  • 讨论开放模式构建以及长期 KB 的整理与质量保证。

实验结果

研究问题

  • RQ1大型、以实体为中心的知识库的基本组成要素与架构是什么?
  • RQ2如何从半结构化与非结构化来源可靠地发现、规范化和增强实体、类型和关系?
  • RQ3哪些开放模式和整理策略支持知识库的长期维护与质量?
  • RQ4从知名知识库案例研究(如 Wikidata、DBpedia、YAGO、OpenK... 等)可以得到哪些经验教训,用于实际的知识库构建与治理?

主要发现

  • 本文综合了构建全面知识库所必需的发现、规范化和增强等一系列技术。
  • 它强调开放世界、按需增长知识库模式,而非固定僵化的模式。
  • 它讨论质量度量、完整性、溯源和生命周期管理,作为知识库维护的核心。
  • 案例研究展示了主要知识库项目如何将原则付诸实践,以及对应用的影响。
  • 该综述将知识库构建与语义检索、问答、NLP 和数据分析等应用相关联,强调跨领域相关性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。