Challenge: Existing emotion corpora collected from twitters and use hashtags are limited in the number of characters.
Approach: They propose to build an emotion corpus based on conversational text data that includes 2.1 million utterances and is partly annotated by ourselves and independent annotators.
Outcome: The proposed corpus includes conversations from movies with more than 2.1 million utterances which are partly annotated by ourselves and independent annotators.

Similar Papers

EmotionLines: An Emotion Corpus of Multi-Party Conversations (L18-1)

Copied to clipboard

Challenge: Emotion is a critical characteristic to distinguish people from machines.
Approach: They propose a dataset with emotions labeling on all utterances in each dialogue . they use Friends TV scripts and Facebook messenger dialogues to collect the data .
Outcome: The proposed dataset is the first with emotions labeling on all utterances in each dialogue based on their textual content.
EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems (2022.lrec-1)

Copied to clipboard

Challenge: Existing emotion-annotated task-oriented corpora are limited in size, label richness, and public availability, creating a bottleneck for downstream tasks.
Approach: They propose a large-scale manually emotion-annotated corpus of task-oriented dialogues based on a multi-domain task-orientated dataset.
Outcome: The proposed method is based on a task-oriented dialogue dataset with 11K dialogues and 83K emotion annotations of user utterances.
Sentence and Clause Level Emotion Annotation, Detection, and Classification in a Multi-Genre Corpus (L18-1)

Copied to clipboard

Challenge: Existing methods for predicting emotion categories are limited due to their multi-label nature . e.g. anger, joy, sadness are difficult to predict due to inherent multi-genre nature - a problem that is often overlooked in single-genrete text.
Approach: They propose to expand existing annotated data to include 8 emotions from Plutchik's Wheel of Emotions . they explore the effectiveness of clause annotation in sentence-level emotion detection and classification .
Outcome: The proposed system is the first to target the clause level and provides emotion classification for movie reviews datasets.
Annotated Corpus for Sentiment Analysis in Odia Language (2020.lrec-1)

Copied to clipboard

Challenge: Existing sentiment analysis models are not available for Odia 1 as it is a resource-poor language.
Approach: They create an annotated Odia corpus and test its usability by training and testing on the corpus using various classifiers.
Outcome: The created corpus contains 2045 Odia sentences from news domain annotated with sentiment labels using a well-defined annotation scheme.
An Emotional Mess! Deciding on a Framework for Building a Dutch Emotion-Annotated Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing frameworks for emotion recognition are limited and do not allow for categorical versus dimensional oppositions.
Approach: They propose to use the emotions joy, love, anger, sadness and fear as well as dimensional models to annotate texts from different domains and topics.
Outcome: The proposed frameworks are well-suited to annotate texts from different domains and topics, but the connotation of the labels strongly depends on the origin of the texts.
A (Psycho-)Linguistically Motivated Scheme for Annotating and Exploring Emotions in a Genre-Diverse Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a linguistic perspective, emotion annotation is considered a difficult task because of the lack of consensus on emotional categories, the fuzziness of boundaries between them or the great variability of emotion expressions types.
Approach: They propose a scheme for emotion annotation and its manual application on a genre-diverse corpus of texts written in french.
Outcome: The proposed method clarifies the main concepts implied by the analysis of emotions as they are expressed in texts and performs a manual annotation campaign on a corpus of 1,594 texts (ca. 515K tokens) of different genres.
An Analysis of Annotated Corpora for Emotion Classification in Text (C18-1)

Copied to clipboard

Challenge: Several datasets have been annotated and published for classification of emotions.
Approach: They aggregated emotion corpora in a common file format with a shared annotation schema . they perform cross-corpus classification experiments to gain insight and a better understanding of differences .
Outcome: The proposed model can be trained on a subset of corpora, but not on all corporata.
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations (2026.findings-acl)

Copied to clipboard

Challenge: Existing datasets face issues such as low quality, limited scale, and incomplete modalities, hindering model performance.
Approach: They propose to use Chinese multimodal datasets to capture authentic emotional interplay from 19 professional actors.
Outcome: The EmotionTalk dataset spans 23.6 hours of dyadic conversations across diverse scenarios.
EmoInHindi: A Multi-label Emotion and Intensity Annotated Dataset in Hindi for Emotion Recognition in Dialogues (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for emotion recognition in dialogues are in English . existing datasets are limited to a few languages like Hindi .
Approach: They propose a large conversational dataset in Hindi for multi-label emotion and intensity recognition in conversations . they use a Wizard-of-Oz manner to annotate dialogues with 16 emotion labels .
Outcome: The proposed dataset contains 1,814 dialogues with 44,247 utterances in Hindi . it is based on a Wizard-of-Oz manner and can detect emotions in conversation .
Who Feels What and Why? Annotation of a Literature Corpus with Semantic Roles of Emotions (C18-1)

Copied to clipboard

Challenge: Emotion analysis and classification is a challenging task which has been tackled with relatively straight-forward approaches.
Approach: They propose to annotate emotion trigger phrases and entities in the roles of experiencers, targets, and causes of the emotion in literature by Project Gutenberg.
Outcome: The proposed corpus supports qualitative literary studies and digital humanities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations