Skip to main content
QUICK REVIEW

[Paper Review] Dynamics of core of language vocabulary

Valery Solovyev, V. V. Bochkarev|arXiv (Cornell University)|May 29, 2017
Language and cultural evolution7 references5 citations
TL;DR

This study analyzes the dynamics of core vocabulary across six European languages using Google Books Ngram data from the past three centuries. It reveals a remarkably stable rate of core word replacement, analogous to the Swadesh list, indicating consistent lexical decay patterns across languages despite differing linguistic and cultural contexts.

ABSTRACT

Studies of the overall structure of vocabulary and its dynamics became possible due to creation of diachronic text corpora, especially Google Books Ngram. This article discusses the question of core change rate and the degree to which the core words cover the texts. Different periods of the last three centuries and six main European languages presented in Google Books Ngram are compared. The main result is high stability of core change rate, which is analogous to stability of the Swadesh list.

Motivation & Objective

  • To investigate the stability and dynamics of core vocabulary across multiple European languages over time.
  • To assess whether the rate of core word replacement remains consistent across different historical periods and languages.
  • To evaluate the extent to which core vocabulary covers general textual content across languages.
  • To compare the dynamics of core vocabulary with established linguistic models such as the Swadesh list.
  • To provide empirical evidence on lexical stability using large-scale diachronic text corpora.

Proposed method

  • Utilizes Google Books Ngram Viewer data covering six major European languages from the 18th to 20th centuries.
  • Identifies core vocabulary by selecting high-frequency, functionally essential words common across languages.
  • Measures the rate of word replacement in the core vocabulary by tracking frequency changes over time.
  • Applies statistical analysis to compare core change rates across different time periods and languages.
  • Employs a comparative framework to assess similarity in core vocabulary dynamics across linguistic families.
  • Validates findings against the Swadesh list as a benchmark for lexical stability.

Experimental results

Research questions

  • RQ1Is the rate of core vocabulary replacement consistent across different European languages over the past three centuries?
  • RQ2How does the stability of core vocabulary compare across distinct historical periods within the same language?
  • RQ3To what extent does the core vocabulary cover the majority of textual content in general corpora?
  • RQ4Does the core word replacement rate exhibit patterns similar to those observed in the Swadesh list?
  • RQ5Are there detectable differences in core vocabulary dynamics between Germanic, Romance, and Slavic language groups?

Key findings

  • The rate of core vocabulary replacement remains highly stable across all six European languages studied, indicating a consistent lexical decay process.
  • Core word change rates are comparable in magnitude to those observed in the Swadesh list, suggesting a universal pattern in lexical evolution.
  • The core vocabulary accounts for a significant and stable proportion of textual content across all time periods and languages.
  • No substantial deviation in core vocabulary dynamics is observed between the 18th, 19th, and 20th centuries, supporting long-term stability.
  • The stability of core change rate is robust across linguistic families, including Germanic, Romance, and Slavic languages.
  • The findings support the hypothesis that core vocabulary evolves at a predictable, near-constant rate, independent of language-specific factors.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.