Challenge: Emotion recognition is an umbrella term for several NLP tasks, but most work on high-resource languages has focused on low-resourced languages.
Approach: They propose to use emotion recognition to describe perceived emotions in 28 different languages and across several domains to identify and annotate the datasets.
Outcome: The proposed datasets cover low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers.

Similar Papers

Representation Mapping: A Novel Approach to Generate High-Quality Multi-Lingual Emotion Lexicons (L18-1)

Copied to clipboard

Challenge: Existing representational frameworks for emotion encoding are incompatible with semantic polarity, resulting in a large amount of incompatible emotion lexicons.
Approach: They propose to map different emotion representation formats onto each other for mutual compatibility and interoperability of language resources.
Outcome: The proposed method produces (near-)gold quality emotion lexicons even in crosslingual settings.
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)

Copied to clipboard

Challenge: a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes .
Approach: They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones .
Outcome: The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions.
An Analysis of Annotated Corpora for Emotion Classification in Text (C18-1)

Copied to clipboard

Challenge: Several datasets have been annotated and published for classification of emotions.
Approach: They aggregated emotion corpora in a common file format with a shared annotation schema . they perform cross-corpus classification experiments to gain insight and a better understanding of differences .
Outcome: The proposed model can be trained on a subset of corpora, but not on all corporata.
Understanding Emotions: A Dataset of Tweets to Study Interactions between Affect Categories (L18-1)

Copied to clipboard

Challenge: a new dataset is used to classify text into positive, negative, and neutral classes . a large amount of work on automatic detecting emotions from text has focused on classifying text into basic emotion categories .
Approach: They use Twitter as the source of the textual data they annotate to find out which emotions often present together in tweets .
Outcome: The proposed dataset is useful for training and testing supervised machine learning algorithms . it is based on the results of the SemEval-2018 task 1: Affect in Tweets .
Learning and Evaluating Emotion Lexicons for 91 Languages (2020.acl-main)

Copied to clipboard

Challenge: Emotion lexicons describe the affective meaning of words but are limited in coverage for most languages.
Approach: They propose a method for creating arbitrarily large emotion lexicons for any target language.
Outcome: The proposed method exceeds human reliability for some languages and variables.
XED: A Multilingual Dataset for Sentiment Analysis and Emotion Detection (2020.coling-main)

Copied to clipboard

Challenge: XED is a multilingual fine-grained emotion dataset for English and other low-resource languages.
Approach: They propose a multilingual fine-grained emotion dataset using Plutchik's Wheel of Emotions and a projection scheme to annotate Finnish and English sentences.
Outcome: The proposed dataset is based on human-annotated Finnish and English sentences and projected annotations for 30 additional languages.
An Emotional Mess! Deciding on a Framework for Building a Dutch Emotion-Annotated Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing frameworks for emotion recognition are limited and do not allow for categorical versus dimensional oppositions.
Approach: They propose to use the emotions joy, love, anger, sadness and fear as well as dimensional models to annotate texts from different domains and topics.
Outcome: The proposed frameworks are well-suited to annotate texts from different domains and topics, but the connotation of the labels strongly depends on the origin of the texts.
A (Psycho-)Linguistically Motivated Scheme for Annotating and Exploring Emotions in a Genre-Diverse Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a linguistic perspective, emotion annotation is considered a difficult task because of the lack of consensus on emotional categories, the fuzziness of boundaries between them or the great variability of emotion expressions types.
Approach: They propose a scheme for emotion annotation and its manual application on a genre-diverse corpus of texts written in french.
Outcome: The proposed method clarifies the main concepts implied by the analysis of emotions as they are expressed in texts and performs a manual annotation campaign on a corpus of 1,594 texts (ca. 515K tokens) of different genres.
Towards Label-Agnostic Emotion Embeddings (2021.emnlp-main)

Copied to clipboard

Challenge: Existing representation schemes for emotion analysis are based on label formats, natural languages, and even disparate model architectures.
Approach: They propose a training scheme that learns a shared latent representation of emotion independent from different label formats, natural languages, and even disparate model architectures.
Outcome: The proposed model performs well on a wide range of datasets without penalizing prediction quality.
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)

Copied to clipboard

Challenge: Language is a powerful means of communication and should be regarded as more than just a collection of tokens.
Approach: They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices.
Outcome: The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations