Challenge: Code-switching (CS) is a common linguistic phenomenon wherein speakers fluidly transition between languages in conversation.
Approach: They propose to use a part-of-speech (POS)-based analysis of Spanish-English and Mandarin-English corpora to examine the propensity of bilinguals to engage in CS.
Outcome: The findings confirm the existence of a statistically significant connection between POS and the likelihood of CS across language pairs, but show that it diminishes as tokens distance themselves from CS instances.

Similar Papers

Discourse-Driven Code-Switching: Analyzing the Role of Content and Communicative Function in Spanish-English Bilingual Speech (2025.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that a range of speaker and listener attributes affect or correlate with the prevalence of code-switching during conversation.
Approach: They analyze the names of entities and dialogue acts present in a Spanish-English spontaneous speech corpus and build a predictive model of CSW.
Outcome: The proposed model is the first to take a discourse-sensitive approach to understanding pragmatic and referential cues of bilingual speech.
Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)

Copied to clipboard

Challenge: Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications.
Approach: They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results .
Outcome: The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions.
Code-Switching and Syntax: A Large-Scale Experiment (2025.findings-acl)

Copied to clipboard

Challenge: Existing theories of code-switching (CS) have been refuted in subsequent investigations.
Approach: They propose to use syntactic information to predict where bilinguals switch languages . they find that syntax alone is sufficient for an automatic system to distinguish between sentences in minimal pairs of CS, to the same degree as bilingual humans.
Outcome: The proposed model can explain why bilinguals switch languages more often than in others, but there is no large-scale, multi-language, cross-phenomena experiment that tests this claim.
A Survey of Code-switching: Linguistic and Social Perspectives for Language Technologies (2021.acl-long)

Copied to clipboard

Challenge: linguistic and social aspects of code-switching are not discussed in the literature in linguistics.
Approach: They propose to examine linguistic and social aspects of code-switching across a wide range of languages in a survey of the literature in linguistics and language technologies.
Outcome: The proposed framework aims to increase the clarity and depth of computational investigations of C-S and bridge the fields so that they might be mutually reinforcing.
Toward the Limitation of Code-Switching in Cross-Lingual Transfer (2022.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown the success of multilingual pretrained models for cross-lingual knowledge transfer.
Approach: They propose to make code-switched sentences replace tokens from multiple languages so they are grammatically consistent . they also consider the similarity between context and the switched tokens to ensure that the newly substituted sentences are grammatically consistent - a limitation that could affect inference .
Outcome: The proposed method outperforms the mBERT and original code-switching method on cross-lingual POS and Named-Entity-Recognition tasks on 30+ languages.
Minimal Pair-Based Evaluation of Code-Switching (2025.acl-long)

Copied to clipboard

Challenge: Existing methods do not have wide language coverage, fail to account for the diverse range of CS phenomena, or do not scale.
Approach: They propose to use minimal pairs of CS to estimate the extent to which large language models (LLMs) use code-switching in the same way as bilinguals.
Outcome: The proposed model assigns higher probability to the naturally occurring CS sentence than to the variant for every language pair.
Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data.
Approach: They investigate the existence of code-switching in the pre-training corpus and categorize it into four types within two quadrants.
Outcome: The proposed approach improves performance across benchmarks and representation space.
End-to-End Speech Translation for Code Switched Speech (2022.findings-acl)

Copied to clipboard

Challenge: Code switching (CS) is the phenomenon of interchangeably using words and phrases from different languages.
Approach: They propose a new ST corpus that extends the joint transcription and translation setup.
Outcome: The proposed model performs well even when no training data is used.
An Integrated Representation of Linguistic and Social Functions of Code-Switching (L18-1)

Copied to clipboard

Challenge: Linguistic studies on code-switching focus on the "how" and "why" of CS . a new model aims to derive CS functions from local and global properties of the code-witched discourse .
Approach: They propose a model that integrates CS phenomena and modalities into a representation that includes local and global properties of the code-switched discourse.
Outcome: The proposed model simplifies the analysis of English/Hindi CS datasets and provides a flexible framework for further studies.
Code-Switched Language Identification is Harder Than You Think (2024.eacl-long)

Copied to clipboard

Challenge: Code switching (CS) is a common phenomenon in written and spoken communication, but is handled poorly by many NLP applications.
Approach: They propose to use CS language identification for corpus building to make it more realistic by scaling it to more languages and considering models with simpler architectures for faster inference.
Outcome: The proposed system is based on a sentence-level multi-label tagging problem and provides recommendations for future work.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations