Code-Switching Patterns Can Be an Effective Route to Improve Performance of Downstream NLP Applications: A Case Study of Humour, Sarcasm and Hate Speech Detection (2020.acl-main)
Copied to clipboard
| Challenge: | In this paper, we demonstrate how code-switching patterns can be utilised to improve various downstream NLP applications. |
| Approach: | They propose to use code-switching patterns to improve various downstream NLP applications. |
| Outcome: | The proposed features can improve humour, sarcasm and hate speech detection tasks. |
Similar Papers
The Decades Progress on Code-Switching Research in NLP: A Systematic Survey on Trends and Challenges (2023.findings-acl)
Copied to clipboard
| Challenge: | Code-Switching is a common phenomenon in written text and conversation . it is not so common to observe code-switching in spoken language and not in written language . |
| Approach: | They present a systematic survey on code-switching research in natural language processing to understand the progress of the past decades and conceptualize the challenges and tasks on the topic. |
| Outcome: | The proposed model combines linguistic theories and machine learning techniques to understand the code-switching phenomenon. |
Towards Code-switched Classification Exploiting Constituent Language Resources (2020.aacl-srw)
Copied to clipboard
| Challenge: | Code-switching is a communicative phenomenon denoting a shift from one language to another within the same speech exchange. |
| Approach: | They propose to convert code-switched data into its constituent high resource languages for use in both monolingual and cross-lingual settings. |
| Outcome: | The proposed code-switching language can be used for multiple downstream tasks . the proposed language increases the F1 score by 22% and 42.5% compared to the state-of-the-art. |
Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)
Copied to clipboard
| Challenge: | Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications. |
| Approach: | They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results . |
| Outcome: | The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions. |
Universal Dependency Parsing for Hindi-English Code-Switching (N18-1)
Copied to clipboard
| Challenge: | Code-switching data often need additional processes such as language identification, normalization and/or back-transliteration to be processed. |
| Approach: | They propose a neural stacking model that leverages part-of-speech tags and syntactic tree annotations in tweets to parse code-switching data. |
| Outcome: | The proposed model is 1.5% better than the augmented model and 3.8% better than one which uses first-best normalization and/or back-transliteration. |
Improving Pretraining Techniques for Code-Switched NLP (2023.acl-long)
Copied to clipboard
| Challenge: | Multilingual pretraining models for code-switched inputs are a key component of NLP applications. |
| Approach: | They propose to use masked language modeling techniques to mask code-switched text that are cognizant of language boundaries prior to masking. |
| Outcome: | The proposed techniques improve performance on two downstream tasks, Question Answering (QA) and Sentiment Analysis (SA), compared to standard pretraining techniques. |
Processing and Understanding Mixed Language Data (D19-2)
Copied to clipboard
| Challenge: | Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text . |
| Approach: | a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities. |
| Outcome: | a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing. |
A Survey of Code-switching: Linguistic and Social Perspectives for Language Technologies (2021.acl-long)
Copied to clipboard
| Challenge: | linguistic and social aspects of code-switching are not discussed in the literature in linguistics. |
| Approach: | They propose to examine linguistic and social aspects of code-switching across a wide range of languages in a survey of the literature in linguistics and language technologies. |
| Outcome: | The proposed framework aims to increase the clarity and depth of computational investigations of C-S and bridge the fields so that they might be mutually reinforcing. |
Improving Code-switched ASR with Linguistic Information (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies on code-switching have been limited to the individual languages, but the results are promising. |
| Approach: | They propose to apply linguistic theories to generate more realistic code-switching text, which is needed for language modelling in ASR. |
| Outcome: | The proposed system improves 2% on English-Spanish code-switching . Equivalence Constraint theory and part-of-speech labelling are particularly helpful for text generation and bring 2% improvement to ASR performance. |
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)
Copied to clipboard
| Challenge: | Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language . |
| Approach: | They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research . |
| Outcome: | This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research . |
Beyond Monolingual Assumptions: A Survey on Code-Switched NLP in the Era of Large Language Models across Modalities (2026.acl-long)
Copied to clipboard
| Challenge: | Amidst the rapid advances of large language models, most LLMs struggle with mixed-language inputs, limited Code-switching datasets, and evaluation biases. |
| Approach: | They propose a roadmap for inclusive datasets, fair evaluation, and linguistically grounded models to achieve truly multilingual intelligence. |
| Outcome: | The proposed frameworks are based on 327 studies spanning five research areas, 15+ NLP tasks, 30+ datasets, and 80+ languages. |