GoEmotions: A Dataset of Fine-Grained Emotions (2020.acl-main)

Copied to clipboard

Challenge: Existing datasets for language-based emotion classification are limited and small . existing datasets lack quality annotations for many different emotion categories .
Approach: They propose to use a large manually annotated dataset to study emotion expressions . they conduct transfer learning experiments with existing emotion benchmarks to test their model .
Outcome: The proposed model achieves an average F1-score of .46, leaving room for improvement.

Similar Papers

Uncovering the Limits of Text-based Emotion Detection (2021.findings-emnlp)

Copied to clipboard

Challenge: Identifying emotions from text is crucial for a variety of downstream tasks.
Approach: They consider the two largest now-available corpora for emotion classification: GoEmotions and Vent.
Outcome: The proposed models outperform the two largest corpora for emotion classification: GoEmotions and Vent.
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts (2026.eacl-long)

Copied to clipboard

Challenge: Recent advances in NLP have greatly improved outcomes in emotion prediction and harmful content detection.
Approach: They propose to classify Vietnamese comments into 27 distinct emotions using a model-based lexical normalization system and a transformer-based model.
Outcome: The proposed corpus of 20,664 social media comments is based on a novel model that can support multiple architectures, but its quality and preprocessing strategies remain key factors influencing performance.
CancerEmo: A Dataset for Fine-Grained Emotion Detection (2020.emnlp-main)

Copied to clipboard

Challenge: a lack of large annotated datasets hinders emotion detection in the health domain . a recent study shows that online sharing of emotions is beneficial to a patient's progress .
Approach: They propose an emotion dataset annotated with eight fine-grained emotions from an online health community.
Outcome: The proposed model achieves an average F1 of 71% on the cancerEmo dataset . the best model achieve a higher F1 than the previous model, which was improved using domain-specific pre-training.
Understanding Emotions: A Dataset of Tweets to Study Interactions between Affect Categories (L18-1)

Copied to clipboard

Challenge: a new dataset is used to classify text into positive, negative, and neutral classes . a large amount of work on automatic detecting emotions from text has focused on classifying text into basic emotion categories .
Approach: They use Twitter as the source of the textual data they annotate to find out which emotions often present together in tweets .
Outcome: The proposed dataset is useful for training and testing supervised machine learning algorithms . it is based on the results of the SemEval-2018 task 1: Affect in Tweets .
Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: a negative emotion is a cognitive bias that affects how we express thoughts and opinions online . a recent study shows that negative words generate more engagement and clicks than positive ones .
Approach: They propose to use readability and linguistic complexity metrics to better understand emotions . they propose to fine-tune three state-of-the-art transformers to detect emotions based on a dataset .
Outcome: The proposed model fails to predict emotions on complex texts, the authors show . they also show that more advanced models fail to predict complex texts .
Hashtags, Emotions, and Comments: A Large-Scale Dataset to Understand Fine-Grained Social Emotions to Online Topics (2020.emnlp-main)

Copied to clipboard

Challenge: A large-scale dataset is collected from Chinese microblog Sina Weibo with over 13 thousand trending topics, emotion votes in 24 fine-grained types from massive participants, and user comments to allow context understanding.
Approach: They use a large-scale dataset from Chinese microblog Sina Weibo to examine readers' responses to online discussion topics.
Outcome: The proposed model outperforms the human model in predicting social emotions in a multilabel classification setting.
A Large-Scale Dataset for Empathetic Response Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Existing empathetic datasets are limited in size and cost due to the cost of manual labor.
Approach: They propose to annotate 1M dialogues with 32 fine-grained emotions and eight empathetic response intents and the Neutral category using a silver dataset.
Outcome: The proposed pipeline compares the quality of the proposed dataset with a state-of-the-art gold dataset using offline experiments and visual validation methods.
EmoNoBa: A Dataset for Analyzing Fine-Grained Emotions on Noisy Bangla Texts (2022.aacl-short)

Copied to clipboard

Challenge: EmoNoBa is a dataset for fine-grained emotion detection on Bangla text . it is based on 22698 comments from social media sites on 12 domains .
Approach: They propose a manually annotated dataset of 22,698 Bangla comments from social media sites on 12 different domains to use for fine-grained emotion detection.
Outcome: The proposed dataset of 22,698 public comments on 12 domains shows that hand-crafted features perform better than neural networks and pre-trained language models.
CHEER-Ekman: Fine-grained Embodied Emotion Classification (2025.acl-short)

Copied to clipboard

Challenge: Emotions manifest through physical experiences and bodily reactions, yet identifying such embodied emotions in text remains understudied.
Approach: They propose to extend existing binary embodied emotion dataset with Ekman’s six basic emotion categories.
Outcome: The proposed dataset outperforms existing methods with large language models.
An Emotional Mess! Deciding on a Framework for Building a Dutch Emotion-Annotated Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing frameworks for emotion recognition are limited and do not allow for categorical versus dimensional oppositions.
Approach: They propose to use the emotions joy, love, anger, sadness and fear as well as dimensional models to annotate texts from different domains and topics.
Outcome: The proposed frameworks are well-suited to annotate texts from different domains and topics, but the connotation of the labels strongly depends on the origin of the texts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations