Skip to main content
QUICK REVIEW

[Paper Review] Phylogeny and geometry of languages from normalized Levenshtein distance

Maurizio Serva|arXiv (Cornell University)|Apr 22, 2011
Language and cultural evolution15 references4 citations
TL;DR

This paper proposes a normalized Levenshtein distance method to objectively quantify lexical distances between languages, avoiding subjective cognate identification. Applied to Malagasy dialects, it reveals a 650 A.D. settlement date and a southeast coastal landing point, supporting oceanic migration routes and providing a reproducible, geometry-based alternative to traditional lexicostatistics.

ABSTRACT

The idea that the distance among pairs of languages can be evaluated from lexical differences seems to have its roots in the work of the French explorer Dumont D'Urville. He collected comparative words lists of various languages during his voyages aboard the Astrolabe from 1826 to 1829 and, in his work about the geographical division of the Pacific, he proposed a method to measure the degree of relation between languages. The method used by the modern lexicostatistics, developed by Morris Swadesh in the 1950s, measures distances from the percentage of shared cognates, which are words with a common historical origin. The weak point of this method is that subjective judgment plays a relevant role. Recently, we have proposed a new automated method which is motivated by the analogy with genetics. The new approach avoids any subjectivity and results can be easily replicated by other scholars. The distance between two languages is defined by considering a renormalized Levenshtein distance between pair of words with the same meaning and averaging on the words contained in a list. The renormalization, which takes into account the length of the words, plays a crucial role, and no sensible results can be found without it. In this paper we give a short review of our automated method and we illustrate it by considering the cluster of Malagasy dialects. We show that it sheds new light on their kinship relation and also that it furnishes a lot of new information concerning the modalities of the settlement of Madagascar.

Motivation & Objective

  • To develop an objective, reproducible method for measuring lexical distances between languages without relying on subjective cognate identification.
  • To overcome the limitations of traditional lexicostatistics, which depend on expert judgment and are biased toward Indo-European language families.
  • To apply a computational, genetics-inspired approach to linguistic data, treating vocabulary as analogous to DNA sequences.
  • To use both phylogenetic and geometric (multidimensional scaling) methods to uncover language relationships and divergence timelines.
  • To test the method on Malagasy dialects to infer migration history, settlement timing, and regional diversity patterns.

Proposed method

  • The method computes a normalized Levenshtein distance between words with the same meaning across languages, defined as d(ω₁,ω₂) = dₗ(ω₁,ω₂)/l(ω₁,ω₂), where dₗ is the standard Levenshtein distance and l is the length of the longer word.
  • Normalization ensures distances are bounded between 0 and 1, making comparisons robust across word lengths and avoiding bias from longer words.
  • Pairwise lexical distances are aggregated into a distance matrix, which is then analyzed using Structural Component Analysis (SCA) to embed languages in a 2D or 3D Euclidean space.
  • SCA preserves both vertical (phylogenetic) and horizontal (contact) relationships, revealing clustering and divergence patterns.
  • Radial distance from the origin in the SCA space is proportional to the time lag from proto-language divergence, enabling dating of linguistic splits.
  • The method is applied to 23 Malagasy dialects and two Austronesian languages (Malay and Maanyan), using Swadesh lists and automated orthographic comparison.

Experimental results

Research questions

  • RQ1What is the objective, reproducible method to measure lexical distances between languages without relying on cognate judgment?
  • RQ2How can linguistic relationships be represented not only as phylogenetic trees but also as geometric configurations in multidimensional space?
  • RQ3When and where did the Malagasy language family originate, based on lexical divergence patterns?
  • RQ4What does the radial distribution of dialects in SCA space reveal about the timing of linguistic divergence?
  • RQ5How do the spatial positions of Malagasy dialects relative to Malay and Maanyan inform migration routes and settlement patterns?

Key findings

  • The method successfully identifies four main dialectal clusters in Malagasy, with the Antandroy variant showing relative isolation but remaining within the same plane as other Malagasy dialects.
  • The radial distribution in the SCA space indicates a divergence lag of approximately 1350 years, corresponding to a founding event around 650 A.D.
  • Dialects such as Antananarivo, Fianarantsoa, Manajary, and Manakara show closer lexical distances to both Maanyan and Malay, suggesting a southeast coastal landing point in Madagascar.
  • The higher linguistic diversity in the southeast region, confirmed by the method, aligns with oceanic current patterns linking Indonesia to that coast.
  • The method supports a migration scenario from Indonesia via the Sunda Strait, with Maanyan as the closest linguistic relative, despite its geographic and cultural distance.
  • The approach enables blind analysis, reducing bias from preconceptions and offering a reproducible alternative to traditional lexicostatistics.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.