[Paper Review] A Survey of Code-switched Speech and Language Processing
A comprehensive survey of code-switching in Speech and NLP, listing datasets, benchmarks, tasks, models, and open challenges in code-switched language processing.
Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches for code-switched Speech and Natural Language Processing. We motivate why processing code-switched text and speech is essential for building intelligent agents and systems that interact with users in multilingual communities. As code-switching data and resources are scarce, we list what is available in various code-switched language pairs with the language processing tasks they can be used for. We review code-switching research in various Speech and NLP applications, including language processing tools and end-to-end systems. We conclude with future directions and open problems in the field.
Motivation & Objective
- Motivate the importance of processing code-switched text and speech for multilingual user interactions.
- Provide a comprehensive catalog of datasets and resources for code-switched language pairs across speech and text tasks.
- Review shared tasks, benchmarks, and evaluation approaches for code-switching in NLP and ASR.
- Summarize modeling approaches and applications, and outline open problems and future directions.
Proposed method
- Summarize linguistic theories of code-switching and translate them into computational considerations for NLP/ASR.
- Catalog available corpora and resources for code-switched speech (ASR/TTS) and text (LID, NER, POS, parsing, QA, NLI, social media data).
- Describe modeling strategies for code-switched systems when data is scarce, including transfer learning, domain adaptation, and synthetic data generation.
- Discuss evaluation benchmarks and methodology for code-switched systems, including language boundaries, matrix language concepts, and cross-language constraints.
- Highlight the role of multilingual models (e.g., multilingual BERT) and embeddings in code-switched NLP.
Experimental results
Research questions
- RQ1What datasets and resources exist for code-switched speech and text across different language pairs?
- RQ2What modeling and evaluation approaches enable effective code-switched ASR and NLP given data scarcity?
- RQ3How have shared tasks and benchmarks shaped progress in code-switched language processing?
- RQ4What open problems and future directions remain for code-switched processing in speech and NLP?
Key findings
- Multiple code-switching datasets exist for speech (e.g., SEAME, HKUST Mandarin-English, CEMOS, CUMIX, MCSM, FACST) and text (LID, NER, POS, parsing, QA, NLI, social media).
- There is a reliance on transfer learning, domain adaptation, and synthetic data to address scarce code-switched resources.
- Shared tasks and benchmarks have driven progress in LID, NER, POS, parsing, QA, and NLI for code-switched data.
- Evaluation of code-switched systems leverages matrix language concepts, language boundary detection, and cross-language constraints as core considerations.
- Large multilingual models and cross-lingual embeddings are explored to handle code-switching in NLP.
- ASR approaches for code-switching include single-pass soft LID decisions, bilingual acoustic models, and data augmentation through synthetic or semi-supervised data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.