Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of tweets annotated for codeswitching and borrowing between Spanish and English is presented . the annotation does not treat common “internet-speak” as codeswitched when used in an otherwise monolingual context. |
| Approach: | They present a new corpus of tweets annotated for codeswitching and borrowing between Spanish and English. |
| Outcome: | The proposed corpus contains 9,500 tweets annotated with codeswitches, borrowings, and named entities. |
Similar Papers
Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling (2022.acl-long)
Copied to clipboard
| Challenge: | a corpus of Spanish newswire rich in unassimilated lexical borrowings is used to identify the language of a word. |
| Approach: | They propose to annotate a corpus of Spanish newswire rich in unassimilated lexical borrowings and evaluate how models perform on this task. |
| Outcome: | The proposed model outperforms models fed with subword embeddings and Transformer-based embeddables on the Spanish newswire corpus. |
Collecting Code-Switched Data from Social Media (L18-1)
Copied to clipboard
| Challenge: | a new method to identify code-switched data from the web is needed . code-witching is defined as the tendency of bilinguals to switch between languages . |
| Approach: | They propose a method that automatically collects code-switched tweets from the web . they use crowd-sourcing to obtain language identifiers for a subset of 8,000 tweets . |
| Outcome: | The proposed method identifies tweets as code-switched in languages L1 and L2 . it is compared to a Spanish-English corpus of code-witched tweets . |
Parsing Tweets into Universal Dependencies (N18-1)
Copied to clipboard
| Challenge: | a new tweet treebank for English is designed to analyze tweets with universal dependencies (UD). |
| Approach: | They extend the universal dependencies guidelines to include special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies. |
| Outcome: | The proposed method outperforms state-of-the-art parsers on other treebanks in accuracy and speed. |
Universal Dependency Parsing for Hindi-English Code-Switching (N18-1)
Copied to clipboard
| Challenge: | Code-switching data often need additional processes such as language identification, normalization and/or back-transliteration to be processed. |
| Approach: | They propose a neural stacking model that leverages part-of-speech tags and syntactic tree annotations in tweets to parse code-switching data. |
| Outcome: | The proposed model is 1.5% better than the augmented model and 3.8% better than one which uses first-best normalization and/or back-transliteration. |
Annotating the Tweebank Corpus on Named Entity Recognition and Building NLP Models for Social Media Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media data such as Twitter messages pose a particular challenge to NLP systems because of their short, noisy nature. |
| Approach: | They create a Twitter-based NER corpus and train Tweet NLP models on it . they annotate named entities in TB2 using Amazon Mechanical Turk . |
| Outcome: | The proposed model outperforms existing models on Twitter and other social media platforms. |
CodeSwitch-Reddit: Exploration of Written Multilingual Discourse in Online Discussion Forums (D19-1)
Copied to clipboard
| Challenge: | a dataset of written code-switched productions is curated from topical threads of multiple bilingual communities on the Reddit discussion platform. |
| Approach: | They analyze a dataset of written code-switched productions curated from multiple bilingual communities on the reddit discussion platform and examine whether findings are carried over to written codeswitching in discussion forums. |
| Outcome: | The proposed dataset can facilitate a range of research and practical activities. |
CodeSwitch-Reddit: Exploration of Written Multilingual Discourse in Online Discussion Forums (D19-55)
Copied to clipboard
| Challenge: | a dataset of written multilingual productions is released to explore the sociolinguistic underpinnings of written code-switching . |
| Approach: | They use a reddit discussion platform to collect written code-switched productions . they examine whether oral code-witching findings are carried over to written code . |
| Outcome: | The proposed dataset can facilitate a range of research and practical activities. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Twitter Universal Dependency Parsing for African-American and Mainstream American English (P18-1)
Copied to clipboard
| Challenge: | We analyze the performance disparities between AAE and Mainstream American English (MAE) because of Twitter-specific conventions and dialectal language. |
| Approach: | They develop a dataset of 500 tweets, 250 of which are in AAE, within the Universal Dependencies 2.0 framework and annotate it. |
| Outcome: | The proposed model improves performance for AAE tweets with no or very little in-domain labeled data and assesses its lexical and syntactic features. |
Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on code-mixing have not been able to model human interactions in context. |
| Approach: | They propose to use a general-purpose code-mixing corpus to model human interactions and relationships in context while maintaining ethical standards. |
| Outcome: | The proposed corpus includes over 355,641 messages spanning various code-mixing patterns, with a primary focus on English, Mandarin, and other languages. |