Māori Loanwords: A Corpus of New Zealand English Tweets (P19-2)

Copied to clipboard

Challenge: Mori loanwords are widely used in New Zealand English for various social functions by New Zealanders within and outside of the Mi community.
Approach: They present a corpus of New Zealand English tweets containing selected Mori words that are likely to be known by New Zealanders who do not speak Mi.
Outcome: The results show that over 30% of the Mori loanwords in tweets are irrelevant . they were manually annotated and used to train machine learning models to filter out irrelevant tweets.

Similar Papers

Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of tweets annotated for codeswitching and borrowing between Spanish and English is presented . the annotation does not treat common “internet-speak” as codeswitched when used in an otherwise monolingual context.
Approach: They present a new corpus of tweets annotated for codeswitching and borrowing between Spanish and English.
Outcome: The proposed corpus contains 9,500 tweets annotated with codeswitches, borrowings, and named entities.
Normalization of Indonesian-English Code-Mixed Twitter Data (D19-55)

Copied to clipboard

Challenge: Twitter is an excellent source of textual data for NLP researches, but it is noisy and often contains typos, slang terms, and non-standard abbreviations.
Approach: They propose a standardization system for Indonesian-English code-mixed Twitter data that includes tokenization, language identification, lexical normalization, and translation.
Outcome: The proposed standardization system is based on four modules for tokenization, language identification, lexical normalization, and translation.
FooTweets: A Bilingual Parallel Corpus of World Cup Tweets (L18-1)

Copied to clipboard

Challenge: a new study analyzes the nature of twitter data and compares it with other social networking websites.
Approach: They develop a parallel corpus of tweets for an English-German pair that can be translated into German using a machine translation tool.
Outcome: The proposed method can be used to translate tweets from English to German using a parallel corpus of 4, 000 tweets.
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)

Copied to clipboard

Challenge: a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages.
Approach: They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction .
Outcome: The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms.
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
Parsing Tweets into Universal Dependencies (N18-1)

Copied to clipboard

Challenge: a new tweet treebank for English is designed to analyze tweets with universal dependencies (UD).
Approach: They extend the universal dependencies guidelines to include special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies.
Outcome: The proposed method outperforms state-of-the-art parsers on other treebanks in accuracy and speed.
A Corpus of Non-Native Written English Annotated for Metaphor (N18-2)

Copied to clipboard

Challenge: Using argumentation-relevant metaphor predicts a holistic score of essay quality, we show .
Approach: They present a corpus of argumentative essays annotated for metaphor by non-native speakers of English . they also examine the relationship between writing proficiency and metaphor use .
Outcome: The proposed corpus is made publicly available and evaluated . it shows that metaphor is a significant predictor of a holistic score of essay quality .
Twitter Universal Dependency Parsing for African-American and Mainstream American English (P18-1)

Copied to clipboard

Challenge: We analyze the performance disparities between AAE and Mainstream American English (MAE) because of Twitter-specific conventions and dialectal language.
Approach: They develop a dataset of 500 tweets, 250 of which are in AAE, within the Universal Dependencies 2.0 framework and annotate it.
Outcome: The proposed model improves performance for AAE tweets with no or very little in-domain labeled data and assesses its lexical and syntactic features.
Classifying the Informative Behaviour of Emoji in Microblogs (L18-1)

Copied to clipboard

Challenge: Emoji are pictographs used in microblogs as emotion markers, but can also represent a wider range of concepts.
Approach: They analyze a corpus of tweets pairs and classify emoji with respect to redundancy . they propose to further investigate the informative behaviour of e-mails using eoji .
Outcome: The proposed model achieved an F-score of 0.7 for emoji use in 2475 tweets pairs.
PoSTWITA-UD: an Italian Twitter Treebank in Universal Dependencies (L18-1)

Copied to clipboard

Challenge: Various approaches and ad hoc resources are needed to provide proper coverage of specific linguistic phenomena.
Approach: They propose to annotate tweets using a well-known dependency-based annotation format . they propose to use the tweets for training NLP systems to improve their performance .
Outcome: The proposed resource can be used for training of NLP systems on social media texts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations