[Paper Review] Multi-Syllable Phonotactic Modelling
This paper presents a novel method for automatically constructing accurate, language-specific phonotactic models using a multi-syllable approach combined with Object-Based Finite-State Modelling (ofs Modelling). By defining multiple syllable classes based on stress and position, and applying a clustering-based generalisation algorithm with a tunable threshold (τ), the approach generates symbolic phonotactic models that closely approximate attested word forms in German, English, and Dutch—significantly reducing overgeneralisation compared to single-syllable models.
This paper describes a novel approach to constructing phonotactic models. The underlying theoretical approach to phonological description is the multisyllable approach in which multiple syllable classes are defined that reflect phonotactically idiosyncratic syllable subcategories. A new finite-state formalism, OFS Modelling, is used as a tool for encoding, automatically constructing and generalising phonotactic descriptions. Language-independent prototype models are constructed which are instantiated on the basis of data sets of phonological strings, and generalised with a clustering algorithm. The resulting approach enables the automatic construction of phonotactic models that encode arbitrarily close approximations of a language's set of attested phonological forms. The approach is applied to the construction of multi-syllable word-level phonotactic models for German, English and Dutch.
Motivation & Objective
- To address the overgeneralisation inherent in single-syllable phonotactic models, which fail to capture position- and stress-dependent phonotactic variation in natural languages.
- To develop a language-independent, data-driven method for constructing symbolic phonotactic models that reflect the true set of attested phonological forms.
- To enable automatic instantiation and generalisation of prototype models using clustering, with control over the degree of generalisation via a tunable threshold (τ).
- To demonstrate that multi-syllable modelling, combined with ofs Modelling, produces more accurate and linguistically plausible phonotactic descriptions than traditional single-syllable approaches.
Proposed method
- The method employs a multi-syllable approach that defines multiple syllable classes based on phonotactic factors such as word position (initial, medial, final) and stress (stressed, unstressed).
- A language-independent prototype model is constructed using Object-Based Finite-State Modelling (ofs Modelling), which encodes all possible combinations of syllable classes and their hierarchical combinations into higher-level constituents.
- The prototype is instantiated with data from phoneme strings extracted from the CELEX lexical database, using fully syllabified, phonetically transcribed forms from German, English, and Dutch.
- A clustering algorithm with a threshold parameter τ is applied to generalise the instantiated models, merging similar syllable classes while preserving linguistically significant distinctions.
- The degree of generalisation is controlled by τ, with lower values producing more clusters and finer distinctions, and higher values leading to broader, more general models.
- The approach enables automatic construction of symbolic phonotactic models that are both highly specific to attested forms and generalisable through controlled clustering.
Experimental results
Research questions
- RQ1Can a multi-syllable approach based on syllable position and stress produce more accurate phonotactic models than single-syllable analyses?
- RQ2How can symbolic phonotactic models be automatically constructed from data without relying on manual, language-specific rule writing?
- RQ3To what extent can a language-independent prototype model be instantiated and generalised using clustering to reflect language-specific phonotactic patterns?
- RQ4How does the choice of threshold τ in the generalisation process affect the balance between model accuracy and coverage?
- RQ5What is the quantitative impact of distinguishing multiple syllable classes on the number of spurious word forms generated by the model?
Key findings
- The multi-syllable approach significantly reduces overgeneralisation: for German, model 1 (single syllable class) encodes 10,598 monosyllabic words, while model 3 (12 syllable classes) encodes only 6,841, matching the actual number in CELEX.
- Model 3, which distinguishes 12 syllable classes, encodes approximately 266 times more bisyllabic forms than the actual number in CELEX, compared to 1.4 times for model 2 and 4.2 times for model 1.
- At τ = 0.3, clustering reveals that stress is a stronger determinant of phonotactic variation than position, with stressed syllables forming distinct clusters across languages.
- In English, both stress and position strongly influence phonotactics, but stress-related differences are more pronounced than position-based ones, as shown by cluster formation at different τ values.
- The method enables the construction of symbolic phonotactic models that are arbitrarily close approximations of a language’s attested phonological forms, with control over generalisation through τ.
- The approach successfully demonstrates that language-independent prototyping with data-driven instantiation and clustering can produce accurate, generalisable, and linguistically meaningful phonotactic models for German, English, and Dutch.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.