[论文解读] Is language evolution grinding to a halt?: Exploring the life and death of words in English fiction.
本研究分析了谷歌图书语料库2012年英语小说子集(1820–2000年)中词汇的出生与消亡速率,采用Jensen-Shannon散度评估不同频率阈值下的波动情况。研究发现,尽管单个词汇的使用频率波动较大,但英语语言的整体统计结构在时间上保持稳定,尽管学术类作品在小说语料库中可能对结果造成偏差。
The Google Books corpus, derived from millions of books in a range of major languages, would seem to offer many possibilities for research into cultural, social, and linguistic evolution. In a previous work, we found that the 2009 and 2012 versions of the unfiltered English data set as well as the 2009 version of the English Fiction data set are all heavily saturated with scientific and medical literature, rendering them unsuitable for rigorous analysis. By contrast, the 2012 version of English Fiction appeared to be uncompromised, and we use this data set to explore language dynamics for English from 1820--2000. We critique a previous method for measuring birth and death rates of words, and provide a robust, principled to examining the volume of word flux across various relative frequency usage thresholds. We use the contributions to the Jensen-Shannon divergence of words crossing thresholds between consecutive decades to illuminate the major driving factors behind the flux. We find that while individual word usage may vary greatly, the overall statistical structure of the language appears to remain fairly stable. We also find indications that scholarly works about fiction are strongly represented in the 2012 English Fiction corpus, and suggest that a future revision of the corpus should attempt to separate critical works from fiction itself.
研究动机与目标
- 评估谷歌图书语料库在研究语言演化方面的可靠性,特别是针对英语小说数据集。
- 通过引入一种基于阈值的系统性方法,解决以往测量词汇出生与消亡速率方法的局限性。
- 探究尽管单个词汇使用存在波动,英语小说语言的统计结构在时间上是否保持稳定。
- 识别2012年英语小说语料库中潜在的偏差,特别是学术类作品混入的影响。
- 建议对语料库进行修订,将评论性文献与原创小说分离,以提升语言研究的数据完整性。
提出的方法
- 使用谷歌图书语料库2012年英语小说子集,因其相对较少受到科学与医学文献的污染,故被选中。
- 采用基于阈值的方法,在不同相对频率区间内测量词汇出生与消亡速率,确保在不同使用水平下均具稳健性。
- 运用Jensen-Shannon散度量化单个词汇在连续十年间对整体词汇波动的贡献。
- 分析跨越频率阈值的词汇如何推动散度变化,识别词汇更替的关键驱动因素。
- 通过检测可能暗示学术内容的异常现象(如与叙事小说无关的高频术语)来验证语料库的完整性。
- 提出一种修订后的语料库结构,将评论性作品与原创小说分离,以减少分析偏差。
实验结果
研究问题
- RQ1谷歌图书语料库2012年英语小说子集在多大程度上仍免受学术与科学文献的污染?
- RQ2词汇出生与消亡速率在不同频率阈值下如何变化,这揭示了词汇动态的哪些特征?
- RQ3跨越频率阈值的单个词汇在推动整体词汇波动方面发挥何种作用,该波动以Jensen-Shannon散度衡量?
- RQ4尽管单个词汇使用存在波动,英语小说语言的统计结构在时间上是否保持稳定?
- RQ5学术类作品关于小说的内容在多大程度上被误标为原创小说,混入2012年英语小说语料库?
主要发现
- 2012年英语小说语料库基本未受科学与医学文献污染,适合用于语言学分析。
- 尽管单个词汇使用存在显著波动,1820至2000年间英语小说语言的整体统计结构仍保持相对稳定。
- 跨越频率阈值的词汇对Jensen-Shannon散度有显著贡献,表明阈值穿越事件是词汇波动的关键驱动力。
- 语料库显示出包含学术类小说评论作品的迹象,可能扭曲词汇频率模式并影响语言推断。
- 本研究识别出未来语料库修订的必要性,即应将评论性文献与原创小说分离,以提升数据质量。
- 所提出的基于阈值的方法为先前词汇出生/消亡速率估算技术提供了一种更系统、更稳健的替代方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。