[Paper Review] The Decades Progress on Code-Switching Research in NLP: A Systematic Survey on Trends and Challenges
This systematic survey synthesizes over 400 NLP papers on code-switching from 1978 to 2022, analyzing trends in languages, tasks, methods, and datasets. It traces the evolution from linguistic theory to neural and pre-trained models, identifies key challenges in multilingual NLP, and outlines future research directions for code-switched language processing in social media and voice applications.
Code-Switching, a common phenomenon in written text and conversation, has been studied over decades by the natural language processing (NLP) research community. Initially, code-switching is intensively explored by leveraging linguistic theories and, currently, more machine-learning oriented approaches to develop models. We introduce a comprehensive systematic survey on code-switching research in natural language processing to understand the progress of the past decades and conceptualize the challenges and tasks on the code-switching topic. Finally, we summarize the trends and findings and conclude with a discussion for future direction and open questions for further investigation.
Motivation & Objective
- To provide a systematic, large-scale survey of code-switching research in NLP across the past decades.
- To analyze the evolution of methods—from rule-based and statistical models to neural and pre-trained models—across language combinations and tasks.
- To identify key research challenges, including data scarcity, multilingual generalization, and task diversity in code-switched NLP.
- To map the landscape of datasets, tasks, and venues to guide future research and highlight underexplored language pairs and applications.
- To examine the influence of linguistic theories on NLP approaches and assess the impact of social media and voice assistants on research trends.
Proposed method
- Collected and manually coded 400+ papers from ACL Anthology, ISCA proceedings, and researcher repositories using standardized keywords like 'code-switching', 'code-mixing', and 'mixed-language'.
- Developed a structured annotation scheme with categories for languages (bilingual, trilingual, 4+), venues (conference, workshop), and methods (rule-based, neural, pre-trained models).
- Mapped NLP tasks into text and speech categories, including sentiment analysis, named entity recognition, speech recognition, and text-to-speech synthesis.
- Tracked publication trends over time using data from *CL and ISCA venues, showing a significant rise in code-switching research post-2015.
- Analyzed the influence of linguistic and socio-linguistic theories on NLP model design, particularly in early and recent work.
- Categorized datasets by source (social media, speech recordings, news, etc.) and evaluated methodological shifts from linguistic constraints to end-to-end deep learning.
Experimental results
Research questions
- RQ1How has the research focus on code-switching evolved in NLP from linguistic theory to machine learning over the past decades?
- RQ2Which language combinations and NLP tasks are most commonly studied, and what are the emerging trends in multilingual NLP?
- RQ3To what extent have linguistic theories influenced the design of modern code-switching NLP models?
- RQ4What are the key methodological shifts in code-switching NLP, from rule-based systems to pre-trained transformer models?
- RQ5What are the most pressing open challenges in code-switching NLP, and what future research directions are most promising?
Key findings
- Code-switching research has grown significantly since 2015, driven by social media and voice-activated devices, with a notable increase in publications at *CL and ISCA venues.
- English-based code-switching (e.g., English-Hindi, English-Tamil) dominates the literature, with 40% of bilingual studies involving English as the second language.
- The shift from rule-based and statistical models to neural and pre-trained models is evident, with 60% of recent papers using transformer-based architectures.
- Despite growth, only 12% of code-switched datasets are multilingual beyond two languages, and 4+ language combinations remain underexplored.
- Speech-related tasks (e.g., ASR, TTS) have seen increasing attention, particularly in South Asian and African language pairs, though data scarcity remains a major bottleneck.
- Only 20% of papers explicitly incorporate linguistic theories into model design, indicating a gap between sociolinguistic insights and NLP methodology.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.