Collecting Code-Switched Data from Social Media (L18-1)

Copied to clipboard

Challenge: a new method to identify code-switched data from the web is needed . code-witching is defined as the tendency of bilinguals to switch between languages .
Approach: They propose a method that automatically collects code-switched tweets from the web . they use crowd-sourcing to obtain language identifiers for a subset of 8,000 tweets .
Outcome: The proposed method identifies tweets as code-switched in languages L1 and L2 . it is compared to a Spanish-English corpus of code-witched tweets .

Similar Papers

TweetTaglish: A Dataset for Investigating Tagalog-English Code-Switching (2022.lrec-1)

Copied to clipboard

Challenge: a large dataset is available to study Tagalog-English code-switching in low-resource settings.
Approach: They propose to use a large dataset to investigate Tagalog-English code-switching . they use linguistic data from Tagalogue and Tagalit-English to investigate their results .
Outcome: The proposed dataset achieves a strong performance benchmark for Tagalog-English code-switching.
Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of tweets annotated for codeswitching and borrowing between Spanish and English is presented . the annotation does not treat common “internet-speak” as codeswitched when used in an otherwise monolingual context.
Approach: They present a new corpus of tweets annotated for codeswitching and borrowing between Spanish and English.
Outcome: The proposed corpus contains 9,500 tweets annotated with codeswitches, borrowings, and named entities.
Universal Dependency Parsing for Hindi-English Code-Switching (N18-1)

Copied to clipboard

Challenge: Code-switching data often need additional processes such as language identification, normalization and/or back-transliteration to be processed.
Approach: They propose a neural stacking model that leverages part-of-speech tags and syntactic tree annotations in tweets to parse code-switching data.
Outcome: The proposed model is 1.5% better than the augmented model and 3.8% better than one which uses first-best normalization and/or back-transliteration.
Lexical Normalization for Code-switched Data and its Effect on POS Tagging (2021.eacl-main)

Copied to clipboard

Challenge: Social media data can be used to improve natural language processing performance, but it is often overlooked by lexical normalization systems.
Approach: They propose three lexical normalization models specifically designed to handle code-switched data and evaluate their performance on POS tags.
Outcome: The proposed models outperform monolingual models and lead to 5.4% performance increase for POS tagging compared to unnormalized input.
Mining Cross-Cultural Differences and Similarities in Social Media (P18-1)

Copied to clipboard

Challenge: a new paper examines the problem of computing cross-cultural differences and similarities in natural language understanding . cross-culture differences are important for cross-lingual research, especially in social media .
Approach: They propose a framework for computing cross-cultural differences and similarities from social media . they propose to use a social media platform to find similar terms for slang across languages .
Outcome: The proposed framework outperforms baseline methods on two novel tasks.
Offensive Content Detection via Synthetic Code-Switched Text (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to detect offensive content in social media platforms are limited by the availability of labeled code-switched data.
Approach: They propose a method for generating synthetic code-switched offensive content data using human-generated data and a keyword classification baseline.
Outcome: The proposed algorithm can be used to generate synthetic code-switched offensive content data and train it on human-generated data.
Detecting Propaganda Techniques in Code-Switched Social Media Text (2023.emnlp-main)

Copied to clipboard

Challenge: a new study aims to detect propaganda in multiple languages using code-switching . social media platforms have made it easier for anyone to spread information to a wide audience .
Approach: They propose to detect propaganda techniques in code-switched texts using a corpus of 1,030 texts . they propose to model multilinguality directly rather than using translation .
Outcome: The proposed method combines different languages within the same text, presenting a challenge for automatic systems.
Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)

Copied to clipboard

Challenge: Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications.
Approach: They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results .
Outcome: The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions.
Processing and Understanding Mixed Language Data (D19-2)

Copied to clipboard

Challenge: Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text .
Approach: a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities.
Outcome: a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing.
Code-Switched Language Identification is Harder Than You Think (2024.eacl-long)

Copied to clipboard

Challenge: Code switching (CS) is a common phenomenon in written and spoken communication, but is handled poorly by many NLP applications.
Approach: They propose to use CS language identification for corpus building to make it more realistic by scaling it to more languages and considering models with simpler architectures for faster inference.
Outcome: The proposed system is based on a sentence-level multi-label tagging problem and provides recommendations for future work.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations