[Paper Review] Are Sounds Sound for Phylogenetic Reconstruction?
The study compares phylogenies inferred from lexical cognates, sound correspondences, and their combination across ten language datasets, finding cognate- and concatenated-based trees are generally closer to gold standard than sound-based trees; Bayesian priors suitable for molecular data can bias language analyses, and sound-based approaches rarely outperform cognate-based ones.
In traditional studies on language evolution, scholars often emphasize the importance of sound laws and sound correspondences for phylogenetic inference of language family trees. However, to date, computational approaches have typically not taken this potential into account. Most computational studies still rely on lexical cognates as major data source for phylogenetic reconstruction in linguistics, although there do exist a few studies in which authors praise the benefits of comparing words at the level of sound sequences. Building on (a) ten diverse datasets from different language families, and (b) state-of-the-art methods for automated cognate and sound correspondence detection, we test, for the first time, the performance of sound-based versus cognate-based approaches to phylogenetic reconstruction. Our results show that phylogenies reconstructed from lexical cognates are topologically closer, by approximately one third with respect to the generalized quartet distance on average, to the gold standard phylogenies than phylogenies reconstructed from sound correspondences.
Motivation & Objective
- Assess whether sound correspondence patterns yield more accurate language phylogenies than lexical cognates.
- Automate cognate detection and sound correspondence inference and evaluate within Bayesian and ML frameworks.
- Cross-validate Bayesian results with maximum likelihood analyses and examine prior effects.
- Compare sound-based phylogenies to cognate-based phylogenies using gold-standard Glottolog trees.
- Investigate whether combining data types provides a superior phylogenetic signal.
Proposed method
- Encode cognate judgments and sound correspondence patterns as binary presence-absence matrices.
- Model evolution with time-reversible binary state CTMC for gain/loss of characters.
- Perform Bayesian phylogenetic inference with MrBayes using standardized priors and examine alpha shape priors.
- Conduct ML phylogenetic inference with RAxML-NG under BIN+G with ML-estimated alpha for rate heterogeneity.
- Trim phonetic alignments to reduce noise and compute sound correspondences with LingRex/LingPy; evaluate with generalized quartet distance (GQD) to Glottolog trees.
- Test three datasets: cognates, sound correspondences, and concatenated matrices.

Experimental results
Research questions
- RQ1Does phylogeny inferred from cognate data better match gold standard trees than phylogeny inferred from sound correspondences?
- RQ2Is a sound-correspondence-based phylogeny ever closer to the gold standard than cognate-based phylogenies?
- RQ3Does concatenating cognate and sound data yield a superior phylogeny compared to using either data type alone?
- RQ4Do Bayesian priors tuned for molecular data bias language phylogenetic results, and can ML analyses reveal such biases?
- RQ5How do sound-based phylogenies compare to cognate-based ones across multiple datasets in terms of generalized quartet distance (GQD)?
Key findings
- Phylogenies from cognate data and from concatenated data are roughly in the same range of accuracy and generally closer to the gold standard than sound-based trees.
- Sound correspondence-based phylogenies never yield the best results in Bayesian analyses across the ten datasets.
- Concatenated data provide the best results for seven of the ten datasets, with cognate data best for the remaining three.
- ML analyses largely corroborate the Bayesian results, showing cognate- and concatenated-based trees are generally closer to the gold standard than sound-based trees.
- Default molecular priors can bias language Bayesian results; complementing with ML analyses helps diagnose such biases.
- The study highlights the need to reassess priors and modeling choices when applying Bayesian methods to linguistic data.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.