[Paper Review] Is language evolution grinding to a halt?: Exploring the life and death of words in English fiction.
This study analyzes word birth and death rates in the 2012 English Fiction subset of the Google Books corpus from 1820 to 2000, using Jensen-Shannon divergence to assess flux across frequency thresholds. It finds that while individual word usage fluctuates, the overall statistical structure of English language remains stable over time, though scholarly works in fiction corpora may bias results.
The Google Books corpus, derived from millions of books in a range of major languages, would seem to offer many possibilities for research into cultural, social, and linguistic evolution. In a previous work, we found that the 2009 and 2012 versions of the unfiltered English data set as well as the 2009 version of the English Fiction data set are all heavily saturated with scientific and medical literature, rendering them unsuitable for rigorous analysis. By contrast, the 2012 version of English Fiction appeared to be uncompromised, and we use this data set to explore language dynamics for English from 1820--2000. We critique a previous method for measuring birth and death rates of words, and provide a robust, principled to examining the volume of word flux across various relative frequency usage thresholds. We use the contributions to the Jensen-Shannon divergence of words crossing thresholds between consecutive decades to illuminate the major driving factors behind the flux. We find that while individual word usage may vary greatly, the overall statistical structure of the language appears to remain fairly stable. We also find indications that scholarly works about fiction are strongly represented in the 2012 English Fiction corpus, and suggest that a future revision of the corpus should attempt to separate critical works from fiction itself.
Motivation & Objective
- To evaluate the reliability of the Google Books corpus for studying linguistic evolution, particularly in the English Fiction dataset.
- To address limitations in prior methods for measuring word birth and death rates by introducing a principled, threshold-based approach.
- To investigate whether the statistical structure of English language remains stable over time despite fluctuations in individual word usage.
- To identify potential biases in the 2012 English Fiction corpus, particularly the inclusion of scholarly works on fiction.
- To recommend corpus revisions that separate critical literature from original fiction to improve data integrity for linguistic research.
Proposed method
- Uses the 2012 English Fiction subset of the Google Books corpus, selected for its relative freedom from scientific and medical literature contamination.
- Applies a threshold-based method to measure word birth and death rates across relative frequency intervals, ensuring robustness across usage levels.
- Employs Jensen-Shannon divergence to quantify the contribution of individual words to overall word flux between consecutive decades.
- Analyzes how words crossing frequency thresholds contribute to divergence, identifying key drivers of lexical change.
- Validates corpus integrity by detecting anomalies suggestive of scholarly content, such as high-frequency terms unrelated to narrative fiction.
- Proposes a revised corpus structure that separates critical works from original fiction to reduce analytical bias.
Experimental results
Research questions
- RQ1To what extent does the 2012 English Fiction corpus of the Google Books dataset remain free from contamination by scholarly and scientific literature?
- RQ2How do word birth and death rates vary across different frequency thresholds, and what does this reveal about lexical dynamics?
- RQ3What role do individual words crossing frequency thresholds play in driving overall word flux, as measured by Jensen-Shannon divergence?
- RQ4Is the statistical structure of the English language in fiction stable over time despite fluctuations in individual word usage?
- RQ5To what extent are scholarly works about fiction misrepresented as original fiction in the 2012 English Fiction corpus?
Key findings
- The 2012 English Fiction corpus is largely uncompromised by scientific and medical literature, making it suitable for linguistic analysis.
- Despite significant variation in individual word usage, the overall statistical structure of the English language in fiction remains fairly stable from 1820 to 2000.
- Words crossing frequency thresholds contribute meaningfully to Jensen-Shannon divergence, indicating that threshold-crossing events are key drivers of lexical flux.
- The corpus shows signs of containing scholarly works about fiction, which may distort word frequency patterns and affect linguistic inferences.
- The study identifies a need for future corpus revisions to separate critical literature from original fiction to improve data quality.
- The proposed threshold-based method provides a more principled and robust alternative to prior word birth/death rate estimation techniques.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.