Challenge: sentiment analysis is used to identify emotional aspects of texts but is limited by its small size and limited range of emotions.
Approach: They propose a Korean sentiment analysis corpus that is limited by its small size and narrow range of emotions . they propose to fine-tune the KOTE dataset and analyze the results for social discrimination .
Outcome: The proposed dataset includes 50,000 Korean online comments, each manually labeled for 43 emotions and NO EMOTION.

Similar Papers

A Japanese Dataset for Subjective and Objective Sentiment Polarity Classification in Micro Blog Domain (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on emotion analysis have studied the analysis of basic emotions and sentiment polarity independently.
Approach: They extend the WRIME dataset with basic emotion intensity from both the writer's subjective and reader's perspective to include the Japanese sentiment polarity.
Outcome: The proposed dataset is the first large-scale corpus to annotate both basic emotions and sentiment polarity labels from both the writer’s and reader’s perspectives.
Understanding Emotions: A Dataset of Tweets to Study Interactions between Affect Categories (L18-1)

Copied to clipboard

Challenge: a new dataset is used to classify text into positive, negative, and neutral classes . a large amount of work on automatic detecting emotions from text has focused on classifying text into basic emotion categories .
Approach: They use Twitter as the source of the textual data they annotate to find out which emotions often present together in tweets .
Outcome: The proposed dataset is useful for training and testing supervised machine learning algorithms . it is based on the results of the SemEval-2018 task 1: Affect in Tweets .
An Analysis of Annotated Corpora for Emotion Classification in Text (C18-1)

Copied to clipboard

Challenge: Several datasets have been annotated and published for classification of emotions.
Approach: They aggregated emotion corpora in a common file format with a shared annotation schema . they perform cross-corpus classification experiments to gain insight and a better understanding of differences .
Outcome: The proposed model can be trained on a subset of corpora, but not on all corporata.
KOAS: Korean Text Offensiveness Analysis System (2021.emnlp-demo)

Copied to clipboard

Challenge: morphological richness and complex syntax of Korean cause difficulties in neural model training.
Approach: They propose a system that exploits contextual and linguistic features and estimates an offensiveness score for a Korean text.
Outcome: The proposed system exploits both contextual and linguistic features and estimates an offensiveness score for a Korean text.
Korean-Specific Emotion Annotation Procedure Using N-Gram-Based Distant Supervision and Korean-Specific-Feature-Based Distant Supervision (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to annotate unlabeled data with emotions are expensive and time-consuming.
Approach: They propose an annotation procedure that leverages Korean emotion lexicons and Korean-specific emotion features to annotate unlabeled data.
Outcome: The proposed procedure compares with the KTEA dataset and a large-scale emotion-labeled dataset.
Hashtags, Emotions, and Comments: A Large-Scale Dataset to Understand Fine-Grained Social Emotions to Online Topics (2020.emnlp-main)

Copied to clipboard

Challenge: A large-scale dataset is collected from Chinese microblog Sina Weibo with over 13 thousand trending topics, emotion votes in 24 fine-grained types from massive participants, and user comments to allow context understanding.
Approach: They use a large-scale dataset from Chinese microblog Sina Weibo to examine readers' responses to online discussion topics.
Outcome: The proposed model outperforms the human model in predicting social emotions in a multilabel classification setting.
Building Large-Scale English and Korean Datasets for Aspect-Level Sentiment Analysis in Automotive Domain (2020.coling-main)

Copied to clipboard

Challenge: Existing datasets in automotive domain cover only three languages due to high cost of human annotation.
Approach: They build large-scale datasets of users’ comments in two languages, English and Korean, for aspect-level sentiment analysis in automotive domain.
Outcome: The datasets consist of 58,000+ commentaspect pairs, which are the largest compared to existing datasets.
K-HATERS: A Hate Speech Detection Corpus in Korean with Target-Specific Ratings (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets on hate speech detection focus on overt forms of hate . however, a majority of these resources are English-centric, focusing on overtones of hate.
Approach: They propose a new corpus for hate speech detection in Korean with target-specific offensiveness ratings that offer a three-point Likert scale.
Outcome: The proposed corpus is the largest offensive language corpus in Korean and offers target-specific ratings on a three-point Likert scale.
GoEmotions: A Dataset of Fine-Grained Emotions (2020.acl-main)

Copied to clipboard

Challenge: Existing datasets for language-based emotion classification are limited and small . existing datasets lack quality annotations for many different emotion categories .
Approach: They propose to use a large manually annotated dataset to study emotion expressions . they conduct transfer learning experiments with existing emotion benchmarks to test their model .
Outcome: The proposed model achieves an average F1-score of .46, leaving room for improvement.
KOLD: Korean Offensive Language Dataset (2022.emnlp-main)

Copied to clipboard

Challenge: Recent directions for offensive language detection focus on English and do not transfer well to other languages because of cultural and linguistic differences.
Approach: They present a Korean offensive language dataset annotated with offensive language comments . they use the comments as training data for Korean BERT and RoBERTa models .
Outcome: The proposed model improves offensiveness detection, target classification, and span detection while having room for improvement for target group classification and span prediction.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations