Analyzing the Role of Part-of-Speech in Code-Switching: A Corpus-Based Study (2024.findings-eacl)
Copied to clipboard
| Challenge: | Code-switching (CS) is a common linguistic phenomenon wherein speakers fluidly transition between languages in conversation. |
| Approach: | They propose to use a part-of-speech (POS)-based analysis of Spanish-English and Mandarin-English corpora to examine the propensity of bilinguals to engage in CS. |
| Outcome: | The findings confirm the existence of a statistically significant connection between POS and the likelihood of CS across language pairs, but show that it diminishes as tokens distance themselves from CS instances. |
Similar Papers
Discourse-Driven Code-Switching: Analyzing the Role of Content and Communicative Function in Spanish-English Bilingual Speech (2025.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has shown that a range of speaker and listener attributes affect or correlate with the prevalence of code-switching during conversation. |
| Approach: | They analyze the names of entities and dialogue acts present in a Spanish-English spontaneous speech corpus and build a predictive model of CSW. |
| Outcome: | The proposed model is the first to take a discourse-sensitive approach to understanding pragmatic and referential cues of bilingual speech. |
Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)
Copied to clipboard
| Challenge: | Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications. |
| Approach: | They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results . |
| Outcome: | The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions. |
Code-Switching and Syntax: A Large-Scale Experiment (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing theories of code-switching (CS) have been refuted in subsequent investigations. |
| Approach: | They propose to use syntactic information to predict where bilinguals switch languages . they find that syntax alone is sufficient for an automatic system to distinguish between sentences in minimal pairs of CS, to the same degree as bilingual humans. |
| Outcome: | The proposed model can explain why bilinguals switch languages more often than in others, but there is no large-scale, multi-language, cross-phenomena experiment that tests this claim. |
A Survey of Code-switching: Linguistic and Social Perspectives for Language Technologies (2021.acl-long)
Copied to clipboard
| Challenge: | linguistic and social aspects of code-switching are not discussed in the literature in linguistics. |
| Approach: | They propose to examine linguistic and social aspects of code-switching across a wide range of languages in a survey of the literature in linguistics and language technologies. |
| Outcome: | The proposed framework aims to increase the clarity and depth of computational investigations of C-S and bridge the fields so that they might be mutually reinforcing. |
Toward the Limitation of Code-Switching in Cross-Lingual Transfer (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have shown the success of multilingual pretrained models for cross-lingual knowledge transfer. |
| Approach: | They propose to make code-switched sentences replace tokens from multiple languages so they are grammatically consistent . they also consider the similarity between context and the switched tokens to ensure that the newly substituted sentences are grammatically consistent - a limitation that could affect inference . |
| Outcome: | The proposed method outperforms the mBERT and original code-switching method on cross-lingual POS and Named-Entity-Recognition tasks on 30+ languages. |
Minimal Pair-Based Evaluation of Code-Switching (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods do not have wide language coverage, fail to account for the diverse range of CS phenomena, or do not scale. |
| Approach: | They propose to use minimal pairs of CS to estimate the extent to which large language models (LLMs) use code-switching in the same way as bilinguals. |
| Outcome: | The proposed model assigns higher probability to the naturally occurring CS sentence than to the variant for every language pair. |
Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training (2025.findings-acl)
Copied to clipboard
Zhijun Wang, Jiahuan Li, Hao Zhou, Rongxiang Weng, Jingang Wang, Xin Huang, Xue Han, Junlan Feng, Chao Deng, Shujian Huang
| Challenge: | Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. |
| Approach: | They investigate the existence of code-switching in the pre-training corpus and categorize it into four types within two quadrants. |
| Outcome: | The proposed approach improves performance across benchmarks and representation space. |
End-to-End Speech Translation for Code Switched Speech (2022.findings-acl)
Copied to clipboard
Orion Weller, Matthias Sperber, Telmo Pires, Hendra Setiawan, Christian Gollan, Dominic Telaar, Matthias Paulik
| Challenge: | Code switching (CS) is the phenomenon of interchangeably using words and phrases from different languages. |
| Approach: | They propose a new ST corpus that extends the joint transcription and translation setup. |
| Outcome: | The proposed model performs well even when no training data is used. |
An Integrated Representation of Linguistic and Social Functions of Code-Switching (L18-1)
Copied to clipboard
| Challenge: | Linguistic studies on code-switching focus on the "how" and "why" of CS . a new model aims to derive CS functions from local and global properties of the code-witched discourse . |
| Approach: | They propose a model that integrates CS phenomena and modalities into a representation that includes local and global properties of the code-switched discourse. |
| Outcome: | The proposed model simplifies the analysis of English/Hindi CS datasets and provides a flexible framework for further studies. |
Code-Switched Language Identification is Harder Than You Think (2024.eacl-long)
Copied to clipboard
| Challenge: | Code switching (CS) is a common phenomenon in written and spoken communication, but is handled poorly by many NLP applications. |
| Approach: | They propose to use CS language identification for corpus building to make it more realistic by scaling it to more languages and considering models with simpler architectures for faster inference. |
| Outcome: | The proposed system is based on a sentence-level multi-label tagging problem and provides recommendations for future work. |