[Paper Review] Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection
UD v2 introduces major guideline updates, expanded multilingual treebanks, and enhanced annotation schemes including morphosyntactic features, relations, and enhanced dependencies across about 90 languages.
Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers. In this paper, we describe version 2 of the guidelines (UD v2), discuss the major changes from UD v1 to UD v2, and give an overview of the currently available treebanks for 90 languages.
Motivation & Objective
- Describe the UD v2 guidelines and the major changes from UD v1 to UD v2.
- Provide an overview of the UD treebank resources available for 90 languages as of UD v2.5.
- Explain the annotation scheme components: tokenization, morphology, syntax, and enhanced dependencies.
- Highlight developments in multilingual parsing and the project’s impact on NLP research.
Proposed method
- Describe tokenization and word segmentation decisions in UD.
- Summarize the universal POS tag set and morphological feature inventory and their extensions in UD v2.
- Explain the UD v2 syntactic relation taxonomy and changes to functional relations, multiword expressions, and coordination.
- Outline the enhanced dependency framework and its five enhancements, including null nodes and propagation of conjuncts.
Experimental results
Research questions
- RQ1What are the key changes from UD v1 to UD v2 in tokenization, morphology, and syntax?
- RQ2How extensive is the UD v2 multilingual resource coverage across languages and treebanks as of v2.5?
- RQ3What are the main design decisions behind UD v2’s annotation scheme and enhanced dependencies?
- RQ4How has UD v2 impacted multilingual parsing research and shared tasks like CoNLL?
- RQ5What is the current scope and diversity of UD treebanks in terms of language families and genres?
Key findings
- UD v2 significantly expands language coverage and treebank resources compared to UD v1, with 157 treebanks and 90 languages by v2.5.
- The annotation scheme maintains a universal POS tag set with extended morphological features and refined syntactic relations, including new or revised categories (e.g., nsubj:pass, obl, cc placement).
- Multiword expressions are revised with a broader use of the compound, fixed, and flat relations, replacing several UD v1 categories and introducing flat:name/flat:foreign subtypes.
- Enhanced dependencies are available for many UD treebanks, enabling explicit representation of implicit relations such as ellipsis, control, and relative clauses, though adoption is partial (24 treebanks).
- The UD project supports multilingual parsing advances and shared tasks, contributing to higher parsing scores and broader language coverage for NLP research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.