Challenge: Existing studies on code-switching have been limited to the individual languages, but the results are promising.
Approach: They propose to apply linguistic theories to generate more realistic code-switching text, which is needed for language modelling in ASR.
Outcome: The proposed system improves 2% on English-Spanish code-switching . Equivalence Constraint theory and part-of-speech labelling are particularly helpful for text generation and bring 2% improvement to ASR performance.

Similar Papers

Language Modeling for Code-Switching: Evaluation, Integration of Monolingual Data, and Discriminative Training (D19-1)

Copied to clipboard

Challenge: Code-switching (CS) is a linguistic phenomenon defined as "the alternation of two languages within a single discourse, sentence or constituent."
Approach: They propose an ASR-motivated evaluation setup which is decoupled from an ASL system and the choice of vocabulary . they propose a discriminative training approach which works better than generative language modeling .
Outcome: The proposed evaluation setup is better than generative language modeling, the authors show . the proposed setup is decoupled from an ASR system and the choice of vocabulary .
Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)

Copied to clipboard

Challenge: Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications.
Approach: They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results .
Outcome: The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions.
New Datasets and Controllable Iterative Data Augmentation Method for Code-switching ASR Error Correction (2023.findings-emnlp)

Copied to clipboard

Challenge: In bilingual or multilingual settings, code-switching ASR has greater challenges and research value.
Approach: They propose a controllable iterative method for improving the performance of mainstream automatic speech recognition systems by using Chinese-English code-switching dialogues.
Outcome: The proposed method achieves the best performance compared with the rule-based, back-translation-based data augmentation methods and large language model ChatGPT.
Improving Language Identification for Code-Switched Speech: The Pivotal Role of Accented English (2026.findings-eacl)

Copied to clipboard

Challenge: Existing models fail to identify English spoken with the accent of the matrix (dominant) language.
Approach: They propose to fine tune existing LID models with accented English to improve code-switched LID . they use a metric that captures relative ranking of identified languages often overlooked by traditional metrics.
Outcome: The proposed model can be fine tuned with small amounts of accented English without degrading performance on monolingual speech.
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models.
Approach: They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages.
Outcome: The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data.
Semi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced Languages (2020.lrec-1)

Copied to clipboard

Challenge: Existing models for code-switching between languages are under-resourced and limited by text and acoustic data.
Approach: They propose to construct four separate bilingual automatic speech recognisers corresponding to four different language pairs between which speakers switch frequently.
Outcome: The proposed models are compared with a non-batch-wise approach and show that they perform better when used with sparse training data.
From English to Code-Switching: Transfer Learning with Strong Morphological Clues (2020.acl-main)

Copied to clipboard

Challenge: Linguistic code-switching (CS) is an understudied area in natural language processing . lack of resources and annotated data makes it difficult to strive for progress in CS-related tasks.
Approach: They propose a method to adapt monolingual models to code-switched text in various tasks . they transfer English knowledge from a pre-trained ELMo model to different code-paired languages .
Outcome: The proposed method outperforms multilingual BERT and homologous CS-unaware models and provides state-of-the-art in CS tasks.
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving (2025.coling-main)

Copied to clipboard

Challenge: More than half of the world's population is presumed to be bilingual . spoken translation of code-switched speech has been under-explored .
Approach: They propose an end-to-end model architecture CoSTA that scaffolds on pretrained ASR and MT modules.
Outcome: The proposed model outperforms existing models by 3.5 BLEU points in spoken translation of code-switched speech.
From Machine Translation to Code-Switching: Generating High-Quality Code-Switched Text (2021.acl-long)

Copied to clipboard

Challenge: a computational model for code-switching text is lacking in the corpus of real text.
Approach: They propose a neural machine translation model to generate Hindi-English code-switched sentences using monolingual Hindi sentences.
Outcome: The proposed model reduces perplexity on a language modeling task and improves on linguistic inference tasks.
End-to-End Speech Translation for Code Switched Speech (2022.findings-acl)

Copied to clipboard

Challenge: Code switching (CS) is the phenomenon of interchangeably using words and phrases from different languages.
Approach: They propose a new ST corpus that extends the joint transcription and translation setup.
Outcome: The proposed model performs well even when no training data is used.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations