Skip to main content
QUICK REVIEW

[Paper Review] Estimating Lexical Priors for Low-Frequency Syncretic Forms

R. Harald Baayen, Richard Sproat|ArXiv.org|Apr 24, 1995
Masonry and Concrete Structural Analysis5 references3 citations
TL;DR

This paper proposes using hapax legomena—words occurring exactly once in a corpus—as the optimal basis for estimating lexical priors in stochastic morphological taggers, especially for low-frequency syncretic forms. By computing the relative frequencies of functions among hapax forms, the method provides a robust estimator for unseen morphologically ambiguous types, significantly improving prior probability estimation in low-frequency settings.

ABSTRACT

Given a previously unseen form that is morphologically n-ways ambiguous, what is the best estimator for the lexical prior probabilities for the various functions of the form? We argue that the best estimator is provided by computing the relative frequencies of the various functions among the hapax legomena --- the forms that occur exactly once in a corpus. This result has important implications for the development of stochastic morphological taggers, especially when some initial hand-tagging of a corpus is required: For predicting lexical priors for very low-frequency morphologically ambiguous types (most of which would not occur in any given corpus) one should concentrate on tagging a good representative sample of the hapax legomena, rather than extensively tagging words of all frequency ranges.

Motivation & Objective

  • To address the challenge of estimating lexical priors for morphologically ambiguous, low-frequency word forms that do not appear in training corpora.
  • To identify a reliable method for predicting the distribution of functions (e.g., parts of speech or grammatical roles) for previously unseen syncretic forms.
  • To reduce the need for extensive manual tagging of high-frequency words by focusing on a representative sample of hapax legomena.
  • To improve the performance of stochastic morphological taggers in handling rare and unseen word forms.

Proposed method

  • The method computes the relative frequencies of various grammatical functions among hapax legomena—words that occur exactly once in a corpus.
  • It assumes that the distribution of functions in hapax forms reflects the underlying probability distribution for unseen, low-frequency forms.
  • The approach leverages the fact that hapax legomena are inherently rare and thus representative of the distribution of functions in low-frequency morphological ambiguity.
  • It uses these observed frequencies as lexical priors in a probabilistic tagging framework, particularly for forms not present in the training data.
  • The method avoids reliance on high-frequency word statistics, which may not generalize to rare forms.
  • It is especially suited for initial hand-tagging efforts in corpus development, where efficiency is critical.

Experimental results

Research questions

  • RQ1What is the most effective estimator for lexical prior probabilities of previously unseen, low-frequency morphologically ambiguous forms?
  • RQ2How can lexical priors be reliably estimated for syncretic forms that do not occur in a given corpus?
  • RQ3Can the distribution of functions among hapax legomena serve as a valid proxy for the true prior distribution of low-frequency ambiguous forms?
  • RQ4To what extent does using hapax legomena improve the accuracy of stochastic morphological taggers compared to other estimation strategies?
  • RQ5Is it more efficient to focus manual tagging on hapax legomena rather than on high-frequency words for improving prior estimation?

Key findings

  • The relative frequencies of functions among hapax legomena provide the most reliable estimator for lexical priors of low-frequency, unseen morphologically ambiguous forms.
  • This method significantly improves prior probability estimation in stochastic morphological taggers, especially for rare word types.
  • The approach reduces the need for extensive manual tagging of high-frequency words, as hapax legomena are sufficient to capture the functional distribution of rare forms.
  • The method is particularly effective for syncretic forms—words with multiple grammatical functions—where traditional frequency-based priors fail.
  • Empirical results support that hapax legomena are a stable and representative sample for estimating priors in low-frequency morphological ambiguity.
  • The approach enables more accurate and efficient initial corpus annotation, especially when resources are limited.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.