[Paper Review] Machine Knowledge: Creation and Curation of Comprehensive Knowledge Bases
A comprehensive survey of methods for automatically constructing and curating large knowledge bases (KBs), covering entity discovery, canonicalization, attribute and relation extraction, open schemas, and long-term KB maintenance with case studies of major KBs.
Equipping machines with comprehensive knowledge of the world's entities and their relationships has been a long-standing goal of AI. Over the last decade, large-scale knowledge bases, also known as knowledge graphs, have been automatically constructed from web contents and text sources, and have become a key asset for search engines. This machine knowledge can be harnessed to semantically interpret textual phrases in news, social media and web tables, and contributes to question answering, natural language processing and data analytics. This article surveys fundamental concepts and practical methods for creating and curating large knowledge bases. It covers models and methods for discovering and canonicalizing entities and their semantic types and organizing them into clean taxonomies. On top of this, the article discusses the automatic extraction of entity-centric properties. To support the long-term life-cycle and the quality assurance of machine knowledge, the article presents methods for constructing open schemas and for knowledge curation. Case studies on academic projects and industrial knowledge graphs complement the survey of concepts and methods.
Motivation & Objective
- Motivate the goal of equipping machines with comprehensive world knowledge for AI applications.
- Survey foundational concepts and architectural design for entity-centric KBs and their lifecycle.
- Present core tasks in KB creation: discovery, canonicalization, augmentation, open schema evolution, and curation.
- Highlight practical methods and design decisions using semi-structured and unstructured sources.
- Provide case studies of prominent KB projects to illustrate principles and challenges.
Proposed method
- Describe knowledge representation foundations for entities, classes, properties, and higher-arity relations.
- Outline design space for KB construction including input sources, quality, and output scope.
- Detail methods for entity discovery and taxonomy construction from diverse sources.
- Explain entity canonicalization including entity linking and matching.
- Present extraction techniques for attributes and relationships from text and semi-structured data.
- Discuss open schema construction and long-term KB curation and quality assurance.
Experimental results
Research questions
- RQ1What are the essential components and architecture of large, entity-centric knowledge bases?
- RQ2How can entities, types, and relations be reliably discovered, canonicalized, and augmented from semi-structured and unstructured sources?
- RQ3What open-schema and curation strategies support long-term maintenance and quality in KBs?
- RQ4What lessons can be drawn from prominent KB CASE studies (e.g., Wikidata, DBpedia, YAGO, OpenK...), for practical KB construction and governance?
Key findings
- The article synthesizes a spectrum of techniques for discovery, canonicalization, and augmentation essential to building comprehensive KBs.
- It emphasizes open-world, pay-as-you-go growth of KB schemas rather than fixed rigid schemas.
- It discusses quality metrics, completeness, provenance, and lifecycle management as central to KB maintenance.
- Case studies illustrate how major KB projects implement principles in practice and their impact on applications.
- The survey links KB construction to applications in semantic search, QA, NLP, and data analytics, highlighting cross-domain relevance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.