[Paper Review] Stability of Syntactic Dialect Classification Over Space and Time
This paper evaluates the spatio-temporal stability of syntax-based dialect classifiers using construction grammar representations across 12 English varieties. It finds that while classification performance decays predictably over time—except in New Zealand English—accuracy varies significantly across locations within countries, revealing that dialect models do not equally represent all local populations, highlighting spatial heterogeneity in grammatical variation.
This paper analyses the degree to which dialect classifiers based on syntactic representations remain stable over space and time. While previous work has shown that the combination of grammar induction and geospatial text classification produces robust dialect models, we do not know what influence both changing grammars and changing populations have on dialect models. This paper constructs a test set for 12 dialects of English that spans three years at monthly intervals with a fixed spatial distribution across 1,120 cities. Syntactic representations are formulated within the usage-based Construction Grammar paradigm (CxG). The decay rate of classification performance for each dialect over time allows us to identify regions undergoing syntactic change. And the distribution of classification accuracy within dialect regions allows us to identify the degree to which the grammar of a dialect is internally heterogeneous. The main contribution of this paper is to show that a rigorous evaluation of dialect classification models can be used to find both variation over space and change over time.
Motivation & Objective
- To assess the temporal and spatial stability of syntax-based dialect classification models using geospatially distributed data.
- To determine whether dialect models remain effective over time despite linguistic and demographic changes.
- To investigate whether dialect models equally represent all populations within a country or show spatial variation in performance.
- To evaluate the role of syntactic constructions in capturing dialectal variation across diverse geographic and temporal contexts.
- To demonstrate that model performance decay and spatial error patterns can reveal linguistic change and internal dialect heterogeneity.
Proposed method
- Constructed a balanced, monthly test set of 1,120 cities across 12 English dialects over a three-year span (2019–2021), maintaining fixed spatial distribution.
- Used construction grammar (CxG) to model syntactic structures, treating each construction as a feature in a bag-of-constructions representation.
- Trained linear SVM classifiers on a fixed 2018 training period and tested performance monthly to assess temporal stability.
- Applied spatial autocorrelation using Empirical Bayes-adjusted Moran’s I to measure spatial structure in classification accuracy across cities.
- Used error rate decay curves to detect temporal changes in dialect models, distinguishing general decay from dialect-specific change.
- Visualized city-level accuracy using maps to identify regions of high and low performance within countries.
Experimental results
Research questions
- RQ1How stable are syntax-based dialect classification models over time across different varieties of English?
- RQ2To what extent does model performance vary across geographic locations within a single dialect region?
- RQ3Can temporal decay rates in classification performance reveal dialect-specific linguistic change?
- RQ4How spatially structured is the variation in model accuracy within a country, and what does this imply about dialect model representativeness?
- RQ5To what degree do dialect models fail to represent certain populations within a dialect area due to internal grammatical heterogeneity?
Key findings
- Most dialects exhibit a consistent decay rate in classification performance over time, indicating general model decay rather than dialect-specific change, except for New Zealand English, which shows detectable temporal change.
- New Zealand English shows the most significant temporal variation, with a distinct change in error distribution over time, suggesting ongoing syntactic change in the dialect.
- Spatial variation in classification accuracy is significant across all countries, with Moran’s I values ranging from 0.17 (Ireland) to 0.70 (Malaysia), indicating strong spatial structure in model performance.
- Within countries, performance varies widely: for example, in New Zealand, accuracy ranges from 8% to 62%, with lower accuracy in rural and geographically distinct regions like Northland and Southland.
- The United States and the United Kingdom show the highest mean accuracy (79% and 73%, respectively), while New Zealand has the lowest (36%), indicating weaker model generalization across its population.
- The study confirms that even the best dialect models do not equally represent all speakers within a dialect, as performance is systematically lower in rural and linguistically distinct regions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.