[论文解读] Lexical semantic change for Ancient Greek and Latin
本文提出并评估了GASC,一种动态贝叶斯混合模型,通过整合分布词表示与体裁元数据,检测古希腊语和拉丁语中的词汇语义变化。结果表明,像GASC这样的贝叶斯模型在检测二元语义变化方面优于最先进词嵌入方法,通过显式建模时间与体裁中的语义演变,实现了高准确率与可解释性。
Change and its precondition, variation, are inherent in languages. Over time, new words enter the lexicon, others become obsolete, and existing words acquire new senses. Associating a word's correct meaning in its historical context is a central challenge in diachronic research. Historical corpora of classical languages, such as Ancient Greek and Latin, typically come with rich metadata, and existing models are limited by their inability to exploit contextual information beyond the document timestamp. While embedding-based methods feature among the current state of the art systems, they are lacking in the interpretative power. In contrast, Bayesian models provide explicit and interpretable representations of semantic change phenomena. In this chapter we build on GASC, a recent computational approach to semantic change based on a dynamic Bayesian mixture model. In this model, the evolution of word senses over time is based not only on distributional information of lexical nature, but also on text genres. We provide a systematic comparison of dynamic Bayesian mixture models for semantic change with state-of-the-art embedding-based models. On top of providing a full description of meaning change over time, we show that Bayesian mixture models are highly competitive approaches to detect binary semantic change in both Ancient Greek and Latin.
研究动机与目标
- 为解决古希腊语和拉丁语等古代语言中计算语义变化模型的空白。
- 探究动态贝叶斯混合模型是否能在古典语料中优于基于嵌入的方法检测语义变化。
- 将体裁元数据整合到语义变化建模中,以提高检测准确率与可解释性。
- 基于专家标注数据,开发用于古代语言中二元语义变化检测的系统性评估框架。
- 证明贝叶斯模型能够嵌入领域知识与历史背景,从而减少因语料不完整带来的偏差。
提出的方法
- 将GASC——一种动态贝叶斯混合模型——适配为联合建模词义概率与体裁随时间的流行度。
- 在贝叶斯框架中整合体裁、作者、风格等类别元数据,以分离语义演变与特定体裁的使用模式。
- 使用狄利克雷过程先验来动态建模词义数量,实现非参数化的语义发现。
- 在具有时间与体裁标注的历史语料上训练模型,估计每个时间周期内词义分配的后验分布。
- 通过精确率、召回率与F1分数,将GASC及其变体SCAN与最先进词嵌入模型(SGNS、TR)及其体裁增强变体进行比较。
- 采用概率框架,通过基于来源与档案历史建模文本存续可能性,以应对文本缺失与语料偏差问题。
实验结果
研究问题
- RQ1动态贝叶斯混合模型是否能在古希腊语和拉丁语中比最先进词嵌入模型更准确地检测二元词汇语义变化?
- RQ2在古典语料中,整合体裁元数据在多大程度上提升了语义变化的检测效果?
- RQ3与黑箱嵌入模型相比,贝叶斯模型在多大程度上提供了可解释且显式的语义演变表征?
- RQ4体裁感知模型是否能比仅依赖分布统计的模型更好地捕捉古代文本中的多义性与语义变异?
- RQ5贝叶斯模型如何扩展以应对历史语言学分析中语料不完整与文本缺失的问题?
主要发现
- 在古希腊语中,GASC在检测二元语义变化方面的F1得分为0.750,优于所有基于嵌入的基线模型,包括SGNS与TR。
- 在拉丁语中,GASC的F1得分为0.788,超过最佳基线(精确率为0.650,召回率为1.000)及所有测试的嵌入模型。
- 该模型在两种语言中均表现出更优性能,其F1得分高于SGNS与TR模型,即使后者经过微调以适应特定体裁情境。
- 体裁感知建模显著提升了检测准确率,尤其在区分多义词如'mus'的技术性与非技术性语义方面表现突出。
- 贝叶斯模型提供了随时间演变的可解释语义轨迹,使研究者能够追踪不同文学体裁中词义的演变过程。
- 本研究证实,贝叶斯模型不仅具有竞争力,而且由于其可解释性以及整合专家知识与历史偏差的能力,更适用于历时研究。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。