| Challenge: | sentiment analysis is used to identify emotional aspects of texts but is limited by its small size and limited range of emotions. |
| Approach: | They propose a Korean sentiment analysis corpus that is limited by its small size and narrow range of emotions . they propose to fine-tune the KOTE dataset and analyze the results for social discrimination . |
| Outcome: | The proposed dataset includes 50,000 Korean online comments, each manually labeled for 43 emotions and NO EMOTION. |
Similar Papers
A Japanese Dataset for Subjective and Objective Sentiment Polarity Classification in Micro Blog Domain (2022.lrec-1)
Copied to clipboard
Haruya Suzuki, Yuto Miyauchi, Kazuki Akiyama, Tomoyuki Kajiwara, Takashi Ninomiya, Noriko Takemura, Yuta Nakashima, Hajime Nagahara
| Challenge: | Existing studies on emotion analysis have studied the analysis of basic emotions and sentiment polarity independently. |
| Approach: | They extend the WRIME dataset with basic emotion intensity from both the writer's subjective and reader's perspective to include the Japanese sentiment polarity. |
| Outcome: | The proposed dataset is the first large-scale corpus to annotate both basic emotions and sentiment polarity labels from both the writer’s and reader’s perspectives. |
Understanding Emotions: A Dataset of Tweets to Study Interactions between Affect Categories (L18-1)
Copied to clipboard
| Challenge: | a new dataset is used to classify text into positive, negative, and neutral classes . a large amount of work on automatic detecting emotions from text has focused on classifying text into basic emotion categories . |
| Approach: | They use Twitter as the source of the textual data they annotate to find out which emotions often present together in tweets . |
| Outcome: | The proposed dataset is useful for training and testing supervised machine learning algorithms . it is based on the results of the SemEval-2018 task 1: Affect in Tweets . |
An Analysis of Annotated Corpora for Emotion Classification in Text (C18-1)
Copied to clipboard
| Challenge: | Several datasets have been annotated and published for classification of emotions. |
| Approach: | They aggregated emotion corpora in a common file format with a shared annotation schema . they perform cross-corpus classification experiments to gain insight and a better understanding of differences . |
| Outcome: | The proposed model can be trained on a subset of corpora, but not on all corporata. |
KOAS: Korean Text Offensiveness Analysis System (2021.emnlp-demo)
Copied to clipboard
San-Hee Park, Kang-Min Kim, Seonhee Cho, Jun-Hyung Park, Hyuntae Park, Hyuna Kim, Seongwon Chung, SangKeun Lee
| Challenge: | morphological richness and complex syntax of Korean cause difficulties in neural model training. |
| Approach: | They propose a system that exploits contextual and linguistic features and estimates an offensiveness score for a Korean text. |
| Outcome: | The proposed system exploits both contextual and linguistic features and estimates an offensiveness score for a Korean text. |
Korean-Specific Emotion Annotation Procedure Using N-Gram-Based Distant Supervision and Korean-Specific-Feature-Based Distant Supervision (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to annotate unlabeled data with emotions are expensive and time-consuming. |
| Approach: | They propose an annotation procedure that leverages Korean emotion lexicons and Korean-specific emotion features to annotate unlabeled data. |
| Outcome: | The proposed procedure compares with the KTEA dataset and a large-scale emotion-labeled dataset. |
Hashtags, Emotions, and Comments: A Large-Scale Dataset to Understand Fine-Grained Social Emotions to Online Topics (2020.emnlp-main)
Copied to clipboard
| Challenge: | A large-scale dataset is collected from Chinese microblog Sina Weibo with over 13 thousand trending topics, emotion votes in 24 fine-grained types from massive participants, and user comments to allow context understanding. |
| Approach: | They use a large-scale dataset from Chinese microblog Sina Weibo to examine readers' responses to online discussion topics. |
| Outcome: | The proposed model outperforms the human model in predicting social emotions in a multilabel classification setting. |
Building Large-Scale English and Korean Datasets for Aspect-Level Sentiment Analysis in Automotive Domain (2020.coling-main)
Copied to clipboard
| Challenge: | Existing datasets in automotive domain cover only three languages due to high cost of human annotation. |
| Approach: | They build large-scale datasets of users’ comments in two languages, English and Korean, for aspect-level sentiment analysis in automotive domain. |
| Outcome: | The datasets consist of 58,000+ commentaspect pairs, which are the largest compared to existing datasets. |
K-HATERS: A Hate Speech Detection Corpus in Korean with Target-Specific Ratings (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing datasets on hate speech detection focus on overt forms of hate . however, a majority of these resources are English-centric, focusing on overtones of hate. |
| Approach: | They propose a new corpus for hate speech detection in Korean with target-specific offensiveness ratings that offer a three-point Likert scale. |
| Outcome: | The proposed corpus is the largest offensive language corpus in Korean and offers target-specific ratings on a three-point Likert scale. |
GoEmotions: A Dataset of Fine-Grained Emotions (2020.acl-main)
Copied to clipboard
| Challenge: | Existing datasets for language-based emotion classification are limited and small . existing datasets lack quality annotations for many different emotion categories . |
| Approach: | They propose to use a large manually annotated dataset to study emotion expressions . they conduct transfer learning experiments with existing emotion benchmarks to test their model . |
| Outcome: | The proposed model achieves an average F1-score of .46, leaving room for improvement. |
KOLD: Korean Offensive Language Dataset (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent directions for offensive language detection focus on English and do not transfer well to other languages because of cultural and linguistic differences. |
| Approach: | They present a Korean offensive language dataset annotated with offensive language comments . they use the comments as training data for Korean BERT and RoBERTa models . |
| Outcome: | The proposed model improves offensiveness detection, target classification, and span detection while having room for improvement for target group classification and span prediction. |