Skip to main content
QUICK REVIEW

[Paper Review] Verb Argument Structure Alternations in Word andSentence Embeddings

Katharina Kann, Alex Warstadt|arXiv (Cornell University)|Jan 1, 2019
Natural Language Processing TechniquesComputer Science37 references30 citations
TL;DR

This paper investigates whether word and sentence embeddings capture fine-grained syntactic distinctions in verb argument structure alternations. Using two new datasets—FAVA (sentence-level) and LaVA (word-level)—the authors test if models can classify acceptable vs. unacceptable verb-frame combinations, finding that performance varies by alternation type and that sentence embeddings lose some information present in word embeddings.

ABSTRACT

Verbs occur in different syntactic environments, or frames. We investigate whether artificial neural networks encode grammatical distinctions necessary for inferring the idiosyncratic frame-selectional properties of verbs. We introduce five datasets, collectively called FAVA, containing in aggregate nearly 10k sentences labeled for grammatical acceptability, illustrating different verbal argument structure alternations. We then test whether models can distinguish acceptable English verb-frame combinations from unacceptable ones using a sentence embedding alone. For converging evidence, we further construct LaVA, a corresponding word-level dataset, and investigate whether the same syntactic features can be extracted from word embeddings. Our models perform reliable classifications for some verbal alternations but not others, suggesting that while these representations do encode fine-grained lexical information, it is incomplete or can be hard to extract. Further, differences between the word- and sentence-level models show that some information present in word embeddings is not passed on to the down-stream sentence embeddings.

Motivation & Objective

  • To investigate whether neural network word and sentence embeddings encode grammatical distinctions needed for verb frame-selectional properties.
  • To evaluate whether sentence embeddings alone can classify acceptable versus unacceptable verb-frame combinations.
  • To compare information retention between word-level and sentence-level embeddings in capturing syntactic features.
  • To identify which types of verb argument structure alternations are more or less predictable from embeddings.

Proposed method

  • Constructed FAVA, a dataset of ~10,000 sentences labeled for grammatical acceptability across five verb argument structure alternations.
  • Built LaVA, a corresponding word-level dataset to enable comparison between word and sentence embeddings.
  • Trained classification models on sentence embeddings to predict acceptability of verb-frame combinations.
  • Trained parallel models on word embeddings to assess whether similar syntactic features are extractable at the word level.
  • Used converging evidence from both word and sentence-level models to evaluate representation quality and information leakage.
  • Evaluated model performance across different alternation types to identify patterns in representational completeness.

Experimental results

Research questions

  • RQ1Can sentence embeddings reliably classify acceptable verb-frame combinations across different argument structure alternations?
  • RQ2To what extent do word embeddings encode syntactic features necessary for frame selection?
  • RQ3Are there differences in classification performance between word-level and sentence-level models for the same alternations?
  • RQ4Which verb argument structure alternations are more effectively captured by neural embeddings?
  • RQ5Is information from word embeddings preserved in downstream sentence embeddings, or is it lost during aggregation?

Key findings

  • Models achieve reliable classification performance for some verb argument structure alternations but not others, indicating incomplete encoding of syntactic distinctions in embeddings.
  • Performance varies significantly across alternation types, suggesting that not all syntactic patterns are equally represented in the embeddings.
  • Word-level models outperform sentence-level models on certain alternations, implying that sentence embeddings may lose or obscure some lexical syntactic information.
  • The study reveals that information present in word embeddings is not always preserved in sentence embeddings, indicating potential information loss during aggregation.
  • Converging evidence from both word and sentence-level models supports the conclusion that while embeddings encode fine-grained lexical information, it is incomplete and context-dependent.
  • The results suggest that current neural network representations are sensitive to specific syntactic patterns but fail to generalize across all verb frame alternations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.