[Paper Review] Corpus-Guided Contrast Sets for Morphosyntactic Feature Detection in Low-Resource English Varieties
This paper presents a human-in-the-loop, corpus-guided method for generating morphosyntactically contrastive training sets to improve automatic detection of linguistic features in low-resource English varieties. By leveraging corpus-driven edits and expert filtering, the approach achieves up to 16-point gains in Prec@100 scores over prior methods on Indian English and African American English, demonstrating improved accuracy and sociolinguistic validity while releasing models and data for reuse.
The study of language variation examines how language varies between and within different groups of speakers, shedding light on how we use language to construct identities and how social contexts affect language use. A common method is to identify instances of a certain linguistic feature - say, the zero copula construction - in a corpus, and analyze the feature's distribution across speakers, topics, and other variables, to either gain a qualitative understanding of the feature's function or systematically measure variation. In this paper, we explore the challenging task of automatic morphosyntactic feature detection in low-resource English varieties. We present a human-in-the-loop approach to generate and filter effective contrast sets via corpus-guided edits. We show that our approach improves feature detection for both Indian English and African American English, demonstrate how it can assist linguistic research, and release our fine-tuned models for use by other researchers.
Motivation & Objective
- To address the challenge of automatic morphosyntactic feature detection in low-resource English varieties, such as Indian English and African American English, where manual annotation is costly and data is scarce.
- To improve the quality and diversity of training data for feature detection by generating contrast sets that are both linguistically meaningful and morphosyntactically similar.
- To validate the method’s effectiveness through quantitative performance metrics and alignment with sociolinguistic findings on age, gender, and regional variation.
- To release fine-tuned models and curated training data for 10 features in Indian English and 17 in African American English to support future research.
Proposed method
- The method generates contrast sets through a corpus-guided edit system that modifies positive examples from a seed set to create semantically and syntactically similar negative examples.
- Human annotators filter and refine the generated contrast sets to ensure linguistic plausibility and morphosyntactic contrastiveness.
- Pretrained language models are fine-tuned on the filtered contrast sets for utterance-level classification of morphosyntactic features.
- The approach combines automatic generation (AutoG, AutoID) with manual filtering (MnlG) and integrates both in a hybrid method (CGEdit) to improve robustness.
- Performance is evaluated using Prec@100, ROC-AUC, and AP metrics across multiple datasets and feature types.
- External validation is performed by comparing model predictions with known sociolinguistic patterns in age, gender, and regional variation.
Experimental results
Research questions
- RQ1Can a corpus-guided, human-in-the-loop method generate high-quality contrast sets that improve morphosyntactic feature detection in low-resource English varieties?
- RQ2How does the performance of the proposed CGEdit method compare to automatic and manual baseline methods in detecting features in Indian English and African American English?
- RQ3To what extent do the model predictions align with established sociolinguistic findings on feature use across age, gender, and regional variation?
- RQ4Can the method generalize across different corpora and time periods, including historical and contemporary data?
Key findings
- The CGEdit method outperforms prior approaches by up to 16 points in Prec@100 scores on Indian English and African American English datasets, demonstrating significant gains in feature detection accuracy.
- On the fwp corpus, the CGEdit method achieved a Prec@100 score of 08.42 for 'non-init. exist. there', compared to 03.29 for AutoG and 00.69 for AutoID, showing substantial improvement.
- The method successfully confirmed sociolinguistic findings: feature frequencies were higher among males and in Princeville compared to DC and Rochester, aligning with prior studies.
- The model predicted higher feature use among younger speakers and lower socioeconomic groups, consistent with findings from Grieser (2019) and Cukor-Avila and Balcazar (2019).
- Standard deviation in feature frequency across age groups was often larger than between age groups, suggesting high individual variability, which the model captured accurately.
- The release of 10 IndE and 17 AAE feature detection models with their training data enables reproducible and extensible research in low-resource language variety analysis.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.