Challenge: a corpus of tweets annotated for codeswitching and borrowing between Spanish and English is presented . the annotation does not treat common “internet-speak” as codeswitched when used in an otherwise monolingual context.
Approach: They present a new corpus of tweets annotated for codeswitching and borrowing between Spanish and English.
Outcome: The proposed corpus contains 9,500 tweets annotated with codeswitches, borrowings, and named entities.

Similar Papers

Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling (2022.acl-long)

Copied to clipboard

Challenge: a corpus of Spanish newswire rich in unassimilated lexical borrowings is used to identify the language of a word.
Approach: They propose to annotate a corpus of Spanish newswire rich in unassimilated lexical borrowings and evaluate how models perform on this task.
Outcome: The proposed model outperforms models fed with subword embeddings and Transformer-based embeddables on the Spanish newswire corpus.
Collecting Code-Switched Data from Social Media (L18-1)

Copied to clipboard

Challenge: a new method to identify code-switched data from the web is needed . code-witching is defined as the tendency of bilinguals to switch between languages .
Approach: They propose a method that automatically collects code-switched tweets from the web . they use crowd-sourcing to obtain language identifiers for a subset of 8,000 tweets .
Outcome: The proposed method identifies tweets as code-switched in languages L1 and L2 . it is compared to a Spanish-English corpus of code-witched tweets .
Parsing Tweets into Universal Dependencies (N18-1)

Copied to clipboard

Challenge: a new tweet treebank for English is designed to analyze tweets with universal dependencies (UD).
Approach: They extend the universal dependencies guidelines to include special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies.
Outcome: The proposed method outperforms state-of-the-art parsers on other treebanks in accuracy and speed.
Universal Dependency Parsing for Hindi-English Code-Switching (N18-1)

Copied to clipboard

Challenge: Code-switching data often need additional processes such as language identification, normalization and/or back-transliteration to be processed.
Approach: They propose a neural stacking model that leverages part-of-speech tags and syntactic tree annotations in tweets to parse code-switching data.
Outcome: The proposed model is 1.5% better than the augmented model and 3.8% better than one which uses first-best normalization and/or back-transliteration.
Annotating the Tweebank Corpus on Named Entity Recognition and Building NLP Models for Social Media Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Social media data such as Twitter messages pose a particular challenge to NLP systems because of their short, noisy nature.
Approach: They create a Twitter-based NER corpus and train Tweet NLP models on it . they annotate named entities in TB2 using Amazon Mechanical Turk .
Outcome: The proposed model outperforms existing models on Twitter and other social media platforms.
CodeSwitch-Reddit: Exploration of Written Multilingual Discourse in Online Discussion Forums (D19-1)

Copied to clipboard

Challenge: a dataset of written code-switched productions is curated from topical threads of multiple bilingual communities on the Reddit discussion platform.
Approach: They analyze a dataset of written code-switched productions curated from multiple bilingual communities on the reddit discussion platform and examine whether findings are carried over to written codeswitching in discussion forums.
Outcome: The proposed dataset can facilitate a range of research and practical activities.
CodeSwitch-Reddit: Exploration of Written Multilingual Discourse in Online Discussion Forums (D19-55)

Copied to clipboard

Challenge: a dataset of written multilingual productions is released to explore the sociolinguistic underpinnings of written code-switching .
Approach: They use a reddit discussion platform to collect written code-switched productions . they examine whether oral code-witching findings are carried over to written code .
Outcome: The proposed dataset can facilitate a range of research and practical activities.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Twitter Universal Dependency Parsing for African-American and Mainstream American English (P18-1)

Copied to clipboard

Challenge: We analyze the performance disparities between AAE and Mainstream American English (MAE) because of Twitter-specific conventions and dialectal language.
Approach: They develop a dataset of 500 tweets, 250 of which are in AAE, within the Universal Dependencies 2.0 framework and annotate it.
Outcome: The proposed model improves performance for AAE tweets with no or very little in-domain labeled data and assesses its lexical and syntactic features.
Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on code-mixing have not been able to model human interactions in context.
Approach: They propose to use a general-purpose code-mixing corpus to model human interactions and relationships in context while maintaining ethical standards.
Outcome: The proposed corpus includes over 355,641 messages spanning various code-mixing patterns, with a primary focus on English, Mandarin, and other languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations