Skip to main content
QUICK REVIEW

[Paper Review] A Survey of Code-switched Speech and Language Processing

Sunayana Sitaram, Khyathi Raghavi Chandu|arXiv (Cornell University)|Mar 25, 2019
Natural Language Processing Techniques212 references53 citations
TL;DR

A comprehensive survey of code-switching in Speech and NLP, listing datasets, benchmarks, tasks, models, and open challenges in code-switched language processing.

ABSTRACT

Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches for code-switched Speech and Natural Language Processing. We motivate why processing code-switched text and speech is essential for building intelligent agents and systems that interact with users in multilingual communities. As code-switching data and resources are scarce, we list what is available in various code-switched language pairs with the language processing tasks they can be used for. We review code-switching research in various Speech and NLP applications, including language processing tools and end-to-end systems. We conclude with future directions and open problems in the field.

Motivation & Objective

  • Motivate the importance of processing code-switched text and speech for multilingual user interactions.
  • Provide a comprehensive catalog of datasets and resources for code-switched language pairs across speech and text tasks.
  • Review shared tasks, benchmarks, and evaluation approaches for code-switching in NLP and ASR.
  • Summarize modeling approaches and applications, and outline open problems and future directions.

Proposed method

  • Summarize linguistic theories of code-switching and translate them into computational considerations for NLP/ASR.
  • Catalog available corpora and resources for code-switched speech (ASR/TTS) and text (LID, NER, POS, parsing, QA, NLI, social media data).
  • Describe modeling strategies for code-switched systems when data is scarce, including transfer learning, domain adaptation, and synthetic data generation.
  • Discuss evaluation benchmarks and methodology for code-switched systems, including language boundaries, matrix language concepts, and cross-language constraints.
  • Highlight the role of multilingual models (e.g., multilingual BERT) and embeddings in code-switched NLP.

Experimental results

Research questions

  • RQ1What datasets and resources exist for code-switched speech and text across different language pairs?
  • RQ2What modeling and evaluation approaches enable effective code-switched ASR and NLP given data scarcity?
  • RQ3How have shared tasks and benchmarks shaped progress in code-switched language processing?
  • RQ4What open problems and future directions remain for code-switched processing in speech and NLP?

Key findings

  • Multiple code-switching datasets exist for speech (e.g., SEAME, HKUST Mandarin-English, CEMOS, CUMIX, MCSM, FACST) and text (LID, NER, POS, parsing, QA, NLI, social media).
  • There is a reliance on transfer learning, domain adaptation, and synthetic data to address scarce code-switched resources.
  • Shared tasks and benchmarks have driven progress in LID, NER, POS, parsing, QA, and NLI for code-switched data.
  • Evaluation of code-switched systems leverages matrix language concepts, language boundary detection, and cross-language constraints as core considerations.
  • Large multilingual models and cross-lingual embeddings are explored to handle code-switching in NLP.
  • ASR approaches for code-switching include single-pass soft LID decisions, bilingual acoustic models, and data augmentation through synthetic or semi-supervised data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.