[Paper Review] A Lexicalized Tree Adjoining Grammar for English
This paper presents a large-scale, lexicalized Tree Adjoining Grammar (TAG) for English, implemented in the XTAG system, which extends standard TAG with feature structures and lexicalization to model complex syntactic phenomena. The grammar covers a broad range of constructions—including passives, wh-clefts, raising, and noun-noun modifications—offering a comprehensive, computationally grounded analysis of English syntax with high precision and extensibility for natural language processing applications.
This document describes a sizable grammar of English written in the TAG formalism and implemented for use with the XTAG system. This report and the grammar described herein supersedes the TAG grammar described in an earlier 1995 XTAG technical report. The English grammar described in this report is based on the TAG formalism which has been extended to include lexicalization, and unification-based feature structures. The range of syntactic phenomena that can be handled is large and includes auxiliaries (including inversion), copula, raising and small clause constructions, topicalization, relative clauses, infinitives, gerunds, passives, adjuncts, it-clefts, wh-clefts, PRO constructions, noun-noun modifications, extraposition, determiner sequences, genitives, negation, noun-verb contractions, sentential adjuncts and imperatives. This technical report corresponds to the XTAG Release 8/31/98. The XTAG grammar is continuously updated with the addition of new analyses and modification of old ones, and an online version of this report can be found at the XTAG web page at http://www.cis.upenn.edu/~xtag/
Motivation & Objective
- To develop a large, computationally grounded grammar of English using the Tree Adjoining Grammar (TAG) formalism.
- To extend TAG with lexicalization and unification-based feature structures to improve syntactic precision and coverage.
- To model a wide range of complex syntactic constructions, including passives, clefts, adjuncts, and small clauses.
- To provide a continuously updated, publicly available grammar resource for linguistic research and NLP systems.
- To supersede earlier versions of the XTAG grammar with more accurate and comprehensive analyses.
Proposed method
- The grammar is built using a lexicalized TAG formalism, where syntactic structures are derived from lexical items and their syntactic features.
- The system employs unification-based feature structures to encode syntactic and semantic constraints, enabling precise control over phrase structure and dependency relations.
- The grammar is implemented in the XTAG system, which supports incremental parsing and linguistic analysis through a modular, rule-based architecture.
- Each syntactic construction is analyzed using tree-adjunct and tree-composition operations, with lexical roots providing the base for syntactic derivation.
- The grammar is manually curated and systematically extended, with updates tracked and published online for community use.
- The formalism supports both constituent and non-constituent structures, allowing for the modeling of discontinuous phenomena like topicalization and extraposition.
Experimental results
Research questions
- RQ1How can a lexicalized TAG formalism be effectively extended to cover a broad range of English syntactic constructions with high precision?
- RQ2What role do feature structures play in enabling fine-grained control over syntactic derivations in a lexicalized TAG framework?
- RQ3How can complex phenomena such as passives, clefts, and raising be systematically encoded within a TAG-based grammar?
- RQ4To what extent can a grammar be modular, extensible, and continuously updated while maintaining consistency and coverage?
- RQ5How does the integration of lexicalization improve the linguistic adequacy and computational efficiency of TAG for natural language processing?
Key findings
- The XTAG grammar successfully models a wide array of syntactic phenomena, including auxiliaries with inversion, copula constructions, and PRO control, with consistent and accurate analyses.
- The integration of unification-based feature structures enables precise handling of syntactic constraints and agreement phenomena across diverse constructions.
- The grammar covers 181 syntactic constructions, as illustrated by 310 pages of text and 181 PostScript figures, demonstrating its comprehensive scope.
- The system supports incremental parsing and linguistic analysis, making it suitable for use in computational linguistics and NLP applications.
- The grammar is continuously updated and available online, with the release corresponding to August 31, 1998, ensuring ongoing relevance and community access.
- The lexicalized approach ensures that each syntactic tree is anchored to a lexical head, improving the linguistic plausibility and interpretability of derivations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.