Skip to main content
QUICK REVIEW

[论文解读] Phylogeny and geometry of languages from normalized Levenshtein distance

Maurizio Serva|arXiv (Cornell University)|Apr 22, 2011
Language and cultural evolution参考文献 15被引用 4
一句话总结

本文提出一种归一化的Levenshtein距离方法,以客观量化语言之间的词汇距离,避免主观的词源识别。该方法应用于马达加斯加方言,揭示了公元650年左右的定居时间与东南沿海的登陆点,支持海洋迁徙路线,并提供了一种可重复的、基于几何的替代传统词汇统计学的方法。

ABSTRACT

The idea that the distance among pairs of languages can be evaluated from lexical differences seems to have its roots in the work of the French explorer Dumont D'Urville. He collected comparative words lists of various languages during his voyages aboard the Astrolabe from 1826 to 1829 and, in his work about the geographical division of the Pacific, he proposed a method to measure the degree of relation between languages. The method used by the modern lexicostatistics, developed by Morris Swadesh in the 1950s, measures distances from the percentage of shared cognates, which are words with a common historical origin. The weak point of this method is that subjective judgment plays a relevant role. Recently, we have proposed a new automated method which is motivated by the analogy with genetics. The new approach avoids any subjectivity and results can be easily replicated by other scholars. The distance between two languages is defined by considering a renormalized Levenshtein distance between pair of words with the same meaning and averaging on the words contained in a list. The renormalization, which takes into account the length of the words, plays a crucial role, and no sensible results can be found without it. In this paper we give a short review of our automated method and we illustrate it by considering the cluster of Malagasy dialects. We show that it sheds new light on their kinship relation and also that it furnishes a lot of new information concerning the modalities of the settlement of Madagascar.

研究动机与目标

  • 开发一种不依赖主观词源识别的客观、可重复的语言间词汇距离测量方法。
  • 克服传统词汇统计学的局限性,后者依赖专家判断,且偏向印欧语系家族。
  • 采用计算方法,借鉴遗传学思路处理语言数据,将词汇视为类似于DNA序列。
  • 结合系统发育分析与几何方法(多维尺度分析),揭示语言关系与分化时间。
  • 将该方法应用于马达加斯加方言,以推断迁徙历史、定居时间与区域多样性模式。

提出的方法

  • 该方法计算跨语言中同一意义词汇之间的归一化Levenshtein距离,定义为 d(ω₁,ω₂) = dₗ(ω₁,ω₂)/l(ω₁,ω₂),其中 dₗ 为标准Levenshtein距离,l 为较长词汇的长度。
  • 归一化确保距离值在0到1之间,使比较在不同词汇长度下保持稳健,避免长词汇带来的偏差。
  • 将成对的词汇距离聚合为距离矩阵,随后通过结构分量分析(SCA)将其嵌入二维或三维欧几里得空间。
  • SCA 同时保留垂直(系统发育)与水平(接触)关系,揭示聚类与分化模式。
  • SCA空间中距原点的径向距离与原始语言分化的时间滞后成正比,从而可对语言分化事件进行定年。
  • 该方法应用于23种马达加斯加方言及两种南岛语系语言(马来语与马亚南语),使用Swadesh词表并结合自动拼写比较。

实验结果

研究问题

  • RQ1是否存在一种不依赖词源判断的客观、可重复的语言间词汇距离测量方法?
  • RQ2语言关系能否不仅以系统发育树的形式,也能以多维空间中的几何构型形式呈现?
  • RQ3基于词汇分化模式,马达加斯加语系家族的起源时间与地点是什么?
  • RQ4SCA空间中方言的径向分布揭示了语言分化的何时发生?
  • RQ5马达加斯加方言相对于马来语与马亚南语的空间位置如何揭示迁徙路线与定居模式?

主要发现

  • 该方法成功识别出马达加斯加方言的四个主要方言群,其中安坦德拉语支表现出相对孤立,但仍位于其他马达加斯加方言所处的同一平面内。
  • SCA空间中的径向分布表明,分化时间滞后约为1350年,对应于公元650年左右的奠基事件。
  • 如安塔那那利佛、菲亚纳兰楚阿、马南加里与马纳卡拉等方言与马亚南语及马来语的词汇距离更近,表明其在马达加斯加东南沿海登陆。
  • 该方法确认东南部地区语言多样性更高,且与连接印度尼西亚与该海岸的洋流模式一致。
  • 该方法支持从印度尼西亚经由苏门答腊海峡的迁徙情景,尽管马亚南语在地理与文化上距离较远,但其为最接近的语言亲属。
  • 该方法支持盲分析,减少先验偏见的影响,为传统词汇统计学提供了一种可重复的替代方案。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。