| Challenge: | Mori loanwords are widely used in New Zealand English for various social functions by New Zealanders within and outside of the Mi community. |
| Approach: | They present a corpus of New Zealand English tweets containing selected Mori words that are likely to be known by New Zealanders who do not speak Mi. |
| Outcome: | The results show that over 30% of the Mori loanwords in tweets are irrelevant . they were manually annotated and used to train machine learning models to filter out irrelevant tweets. |
Similar Papers
Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of tweets annotated for codeswitching and borrowing between Spanish and English is presented . the annotation does not treat common “internet-speak” as codeswitched when used in an otherwise monolingual context. |
| Approach: | They present a new corpus of tweets annotated for codeswitching and borrowing between Spanish and English. |
| Outcome: | The proposed corpus contains 9,500 tweets annotated with codeswitches, borrowings, and named entities. |
Normalization of Indonesian-English Code-Mixed Twitter Data (D19-55)
Copied to clipboard
| Challenge: | Twitter is an excellent source of textual data for NLP researches, but it is noisy and often contains typos, slang terms, and non-standard abbreviations. |
| Approach: | They propose a standardization system for Indonesian-English code-mixed Twitter data that includes tokenization, language identification, lexical normalization, and translation. |
| Outcome: | The proposed standardization system is based on four modules for tokenization, language identification, lexical normalization, and translation. |
FooTweets: A Bilingual Parallel Corpus of World Cup Tweets (L18-1)
Copied to clipboard
| Challenge: | a new study analyzes the nature of twitter data and compares it with other social networking websites. |
| Approach: | They develop a parallel corpus of tweets for an English-German pair that can be translated into German using a machine translation tool. |
| Outcome: | The proposed method can be used to translate tweets from English to German using a parallel corpus of 4, 000 tweets. |
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)
Copied to clipboard
| Challenge: | a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages. |
| Approach: | They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction . |
| Outcome: | The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
Parsing Tweets into Universal Dependencies (N18-1)
Copied to clipboard
| Challenge: | a new tweet treebank for English is designed to analyze tweets with universal dependencies (UD). |
| Approach: | They extend the universal dependencies guidelines to include special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies. |
| Outcome: | The proposed method outperforms state-of-the-art parsers on other treebanks in accuracy and speed. |
A Corpus of Non-Native Written English Annotated for Metaphor (N18-2)
Copied to clipboard
| Challenge: | Using argumentation-relevant metaphor predicts a holistic score of essay quality, we show . |
| Approach: | They present a corpus of argumentative essays annotated for metaphor by non-native speakers of English . they also examine the relationship between writing proficiency and metaphor use . |
| Outcome: | The proposed corpus is made publicly available and evaluated . it shows that metaphor is a significant predictor of a holistic score of essay quality . |
Twitter Universal Dependency Parsing for African-American and Mainstream American English (P18-1)
Copied to clipboard
| Challenge: | We analyze the performance disparities between AAE and Mainstream American English (MAE) because of Twitter-specific conventions and dialectal language. |
| Approach: | They develop a dataset of 500 tweets, 250 of which are in AAE, within the Universal Dependencies 2.0 framework and annotate it. |
| Outcome: | The proposed model improves performance for AAE tweets with no or very little in-domain labeled data and assesses its lexical and syntactic features. |
Classifying the Informative Behaviour of Emoji in Microblogs (L18-1)
Copied to clipboard
| Challenge: | Emoji are pictographs used in microblogs as emotion markers, but can also represent a wider range of concepts. |
| Approach: | They analyze a corpus of tweets pairs and classify emoji with respect to redundancy . they propose to further investigate the informative behaviour of e-mails using eoji . |
| Outcome: | The proposed model achieved an F-score of 0.7 for emoji use in 2475 tweets pairs. |
PoSTWITA-UD: an Italian Twitter Treebank in Universal Dependencies (L18-1)
Copied to clipboard
Manuela Sanguinetti, Cristina Bosco, Alberto Lavelli, Alessandro Mazzei, Oronzo Antonelli, Fabio Tamburini
| Challenge: | Various approaches and ad hoc resources are needed to provide proper coverage of specific linguistic phenomena. |
| Approach: | They propose to annotate tweets using a well-known dependency-based annotation format . they propose to use the tweets for training NLP systems to improve their performance . |
| Outcome: | The proposed resource can be used for training of NLP systems on social media texts. |