Creation of Corpus and analysis in Code-Mixed Kannada-English Twitter data for Emotion Prediction (2020.coling-main)
Copied to clipboard
| Challenge: | Existing work on emotion prediction for resource-rich languages has focused on code-mixed social media corpus but not on Kannada-English code-mixed Twitter data. |
| Approach: | They analyze Kannada-English code-mixed Twitter corpus annotated with their respective ‘Emotion’ for each tweet. |
| Outcome: | The proposed model based on Kannada-English code-mixed Twitter corpus yielded an accuracy of 30% and 32% respectively. |
Similar Papers
Corpus Creation and Emotion Prediction for Hindi-English Code-Mixed Social Media Text (N18-4)
Copied to clipboard
| Challenge: | Emotion Prediction is a natural language processing task dealing with detection and classification of emotions in monolingual and bilingual texts. |
| Approach: | They propose a machine learning system which uses various machine learning techniques to detect emotion associated with tweets. |
| Outcome: | The proposed system uses various machine learning techniques to detect emotion associated with the text. |
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)
Copied to clipboard
| Challenge: | a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages. |
| Approach: | They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction . |
| Outcome: | The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms. |
Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)
Copied to clipboard
| Challenge: | a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people. |
| Approach: | They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook. |
| Outcome: | The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field. |
EmoEvent: A Multilingual Emotion Corpus based on different Events (2020.lrec-1)
Copied to clipboard
| Challenge: | In recent years, emotion detection in text has become more popular due to its potential applications in fields such as psychology, marketing, political science, among others. |
| Approach: | They propose to use an annotated dataset to identify emotions in tweets from different events that took place in April 2019 to validate the effectiveness of the data set. |
| Outcome: | The proposed method is based on a multilingual emotion data set based in different events that took place in April 2019 in English and Spanish. |
Normalization of Indonesian-English Code-Mixed Twitter Data (D19-55)
Copied to clipboard
| Challenge: | Twitter is an excellent source of textual data for NLP researches, but it is noisy and often contains typos, slang terms, and non-standard abbreviations. |
| Approach: | They propose a standardization system for Indonesian-English code-mixed Twitter data that includes tokenization, language identification, lexical normalization, and translation. |
| Outcome: | The proposed standardization system is based on four modules for tokenization, language identification, lexical normalization, and translation. |
Corpus Creation and Analysis for Named Entity Recognition in Telugu-English Code-Mixed Social Media Data (P19-2)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a subtask of Information Extraction in NLP. |
| Approach: | They present a Telugu-English code-mixed corpus with the corresponding named entity tags. |
| Outcome: | The proposed model scored 0.96, 0.94 and 0.95 on a Telugu-English code-mixed corpus. |
MojiTalk: Generating Emotional Responses at Scale (P18-1)
Copied to clipboard
| Challenge: | Existing studies on emotion-generating systems focus on small sets of labeled datasets. |
| Approach: | They propose to leverage Twitter data that are naturally labeled with emojis to generate emotional responses. |
| Outcome: | The proposed models can generate high-quality conversation responses in accordance with designated emotions. |
Semi-Automatic Construction and Refinement of an Annotated Corpus for a Deep Learning Framework for Emotion Classification (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for emotion classification are expensive and require a large corpus of data. |
| Approach: | They propose a method for creating a semi-automatically constructed emotion corpus by correcting errors in the corpus. |
| Outcome: | The proposed method improves the quality of the emotion labels by correcting errors. |
Multi-domain Tweet Corpora for Sentiment Analysis: Resource Creation and Evaluation (2020.lrec-1)
Copied to clipboard
| Challenge: | a huge amount of content is being generated every day due to the pervasiveness of social media. |
| Approach: | They firstly create a multi-domain tweet sentiment corpora and then establish a deep neural network based baseline framework to address the above mentioned issues. |
| Outcome: | The proposed dataset achieves 84.65% accuracy for sentiment analysis using a neural network, long short term memory, and gated recurrent unit (GRU). |
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. |
| Approach: | They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors. |
| Outcome: | The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups. |