[Paper Review] What do complexity measures measure? Correlating and validating corpus-based measures of morphological complexity
This study evaluates eight corpus-based measures of morphological complexity across 30 languages, finding strong correlations and a dominant single underlying dimension through principal component analysis (92.6% variance explained). Despite theoretical distinctions between enumerative and integrative complexity, the measures primarily reflect one coherent dimension of morphological complexity, with combined use offering more reliable results than single metrics.
We present an analysis of eight measures used for quantifying morphological complexity of natural languages. The measures we study are corpus-based measures of morphological complexity with varying requirements for corpus annotation. We present similarities and differences between these measures visually and through correlation analyses, as well as their relation to the relevant typological variables. Our analysis focuses on whether these `measures' are measures of the same underlying variable, or whether they measure more than one dimension of morphological complexity. The principal component analysis indicates that the first principal component explains 92.62 % of the variation in eight measures, indicating a strong linear dependence between the complexity measures studied.
Motivation & Objective
- To assess whether widely used corpus-based morphological complexity measures quantify the same underlying linguistic construct.
- To investigate whether these measures capture multiple dimensions of complexity, such as enumerative (number of morphosyntactic distinctions) and integrative (predictability of forms).
- To validate the measures against typological features from the WALS database and inflection accuracy data.
- To determine whether combining multiple measures improves reliability in assessing morphological complexity.
- To examine the impact of corpus size and irregularity on model performance and complexity estimation.
Proposed method
- Applied principal component analysis (PCA) to eight morphological complexity measures across 30 languages to identify underlying dimensions of variation.
- Used corpus-derived inflection tables from multiple treebanks (e.g., UD, TüBa-D/Z) to compute complexity scores, ensuring consistency with real language use.
- Correlated the complexity measures with WALS typological features and negative inflection accuracy scores to assess external validity.
- Performed correlation analysis between all pairs of complexity measures to evaluate their interdependence.
- Analyzed the impact of corpus size and frequency of irregular forms on model performance and complexity estimation.
- Visualized language rankings across measures and PCA dimensions to assess clustering by language family and typological similarity.
Experimental results
Research questions
- RQ1Do the eight corpus-based morphological complexity measures assess the same underlying linguistic construct?
- RQ2To what extent do these measures reflect multiple dimensions of morphological complexity, such as enumerative and integrative complexity?
- RQ3How do the complexity measures correlate with external linguistic variables like WALS typological features and inflection accuracy?
- RQ4Why do complexity measures show a positive correlation with negative inflection accuracy, contrary to expectations of a negative correlation?
- RQ5Can combining multiple complexity measures yield a more stable and reliable estimate of morphological complexity than using a single measure?
Key findings
- The first principal component explains 92.6213% of the variance in the eight complexity measures, indicating strong linear dependence and a dominant single underlying dimension.
- Despite theoretical distinctions between enumerative and integrative complexity, the measures primarily reflect one coherent dimension of morphological complexity.
- Measures sensitive to derivational morphology and compounding (e.g., WH and LH) yield higher scores for languages like Vietnamese and English, which are otherwise ranked lower in complexity.
- All complexity measures show positive correlations with each other, supporting their shared focus on a common construct.
- The measure MSP shows the highest correlation with negative inflection accuracy (0.2837), suggesting it is most sensitive to irregularity and processing difficulty.
- Treebanks from the same or closely related languages cluster closely in the PCA space, indicating that the measures reliably reflect typological and genetic relationships.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.