| Challenge: | a new method to identify code-switched data from the web is needed . code-witching is defined as the tendency of bilinguals to switch between languages . |
| Approach: | They propose a method that automatically collects code-switched tweets from the web . they use crowd-sourcing to obtain language identifiers for a subset of 8,000 tweets . |
| Outcome: | The proposed method identifies tweets as code-switched in languages L1 and L2 . it is compared to a Spanish-English corpus of code-witched tweets . |
Similar Papers
TweetTaglish: A Dataset for Investigating Tagalog-English Code-Switching (2022.lrec-1)
Copied to clipboard
| Challenge: | a large dataset is available to study Tagalog-English code-switching in low-resource settings. |
| Approach: | They propose to use a large dataset to investigate Tagalog-English code-switching . they use linguistic data from Tagalogue and Tagalit-English to investigate their results . |
| Outcome: | The proposed dataset achieves a strong performance benchmark for Tagalog-English code-switching. |
Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of tweets annotated for codeswitching and borrowing between Spanish and English is presented . the annotation does not treat common “internet-speak” as codeswitched when used in an otherwise monolingual context. |
| Approach: | They present a new corpus of tweets annotated for codeswitching and borrowing between Spanish and English. |
| Outcome: | The proposed corpus contains 9,500 tweets annotated with codeswitches, borrowings, and named entities. |
Universal Dependency Parsing for Hindi-English Code-Switching (N18-1)
Copied to clipboard
| Challenge: | Code-switching data often need additional processes such as language identification, normalization and/or back-transliteration to be processed. |
| Approach: | They propose a neural stacking model that leverages part-of-speech tags and syntactic tree annotations in tweets to parse code-switching data. |
| Outcome: | The proposed model is 1.5% better than the augmented model and 3.8% better than one which uses first-best normalization and/or back-transliteration. |
Lexical Normalization for Code-switched Data and its Effect on POS Tagging (2021.eacl-main)
Copied to clipboard
| Challenge: | Social media data can be used to improve natural language processing performance, but it is often overlooked by lexical normalization systems. |
| Approach: | They propose three lexical normalization models specifically designed to handle code-switched data and evaluate their performance on POS tags. |
| Outcome: | The proposed models outperform monolingual models and lead to 5.4% performance increase for POS tagging compared to unnormalized input. |
Mining Cross-Cultural Differences and Similarities in Social Media (P18-1)
Copied to clipboard
| Challenge: | a new paper examines the problem of computing cross-cultural differences and similarities in natural language understanding . cross-culture differences are important for cross-lingual research, especially in social media . |
| Approach: | They propose a framework for computing cross-cultural differences and similarities from social media . they propose to use a social media platform to find similar terms for slang across languages . |
| Outcome: | The proposed framework outperforms baseline methods on two novel tasks. |
Offensive Content Detection via Synthetic Code-Switched Text (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods to detect offensive content in social media platforms are limited by the availability of labeled code-switched data. |
| Approach: | They propose a method for generating synthetic code-switched offensive content data using human-generated data and a keyword classification baseline. |
| Outcome: | The proposed algorithm can be used to generate synthetic code-switched offensive content data and train it on human-generated data. |
Detecting Propaganda Techniques in Code-Switched Social Media Text (2023.emnlp-main)
Copied to clipboard
| Challenge: | a new study aims to detect propaganda in multiple languages using code-switching . social media platforms have made it easier for anyone to spread information to a wide audience . |
| Approach: | They propose to detect propaganda techniques in code-switched texts using a corpus of 1,030 texts . they propose to model multilinguality directly rather than using translation . |
| Outcome: | The proposed method combines different languages within the same text, presenting a challenge for automatic systems. |
Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)
Copied to clipboard
| Challenge: | Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications. |
| Approach: | They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results . |
| Outcome: | The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions. |
Processing and Understanding Mixed Language Data (D19-2)
Copied to clipboard
| Challenge: | Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text . |
| Approach: | a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities. |
| Outcome: | a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing. |
Code-Switched Language Identification is Harder Than You Think (2024.eacl-long)
Copied to clipboard
| Challenge: | Code switching (CS) is a common phenomenon in written and spoken communication, but is handled poorly by many NLP applications. |
| Approach: | They propose to use CS language identification for corpus building to make it more realistic by scaling it to more languages and considering models with simpler architectures for faster inference. |
| Outcome: | The proposed system is based on a sentence-level multi-label tagging problem and provides recommendations for future work. |