Challenge: Existing databases for emotion recognition are limited due to privacy and legal issues.
Approach: They propose to collect emotional speech data from actors and non-actors using designed drama situations and annotate them manually using a hybrid strategy.
Outcome: The proposed database is based on simulated parallel, semi-natural, and (near to) natural databases.

Similar Papers

A Comparative Cross Language View On Acted Databases Portraying Basic Emotions Utilising Machine Learning (2022.lrec-1)

Copied to clipboard

Challenge: Since several decades emotional databases have been recorded by various laboratories.
Approach: They propose to model similarity as performance in cross database machine learning experiments and to analyze a manually picked set of four acoustic features that represent different phonetic areas.
Outcome: The proposed sets of features represent different phonetic areas and are comparable across languages.
A Dataset for Speech Emotion Recognition in Greek Theatrical Plays (2022.lrec-1)

Copied to clipboard

Challenge: Speech Emotion Recognition (SER) is a task that is difficult to perform by humans due to subjectiveness of the emotional content.
Approach: They propose to use GreThE to collect data for speech emotion recognition in Greek plays.
Outcome: The proposed dataset contains utterances from various actors and plays, along with respective valence and arousal annotations.
M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database (2022.acl-long)

Copied to clipboard

Challenge: Existing data resources to support multimodal affective analysis in dialogues are limited in scale and diversity.
Approach: They propose a multimodal multi-scene multi-label Emotional Dialogue dataset, M3ED, which contains 990 dyadic emotional dialogues from 56 different TV series.
Outcome: The proposed dataset contains 990 dyadic emotional dialogues from 56 different TV series, a total of 9,082 turns and 24,449 utterances.
MELD-ST: An Emotion-aware Speech Translation Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Emotion plays a crucial role in human conversation.
Approach: They present a MELD-ST dataset for the emotion-aware speech translation task . they show that fine-tuning with emotion labels can enhance translation performance .
Outcome: The proposed dataset shows that fine tuning with emotion labels can improve translation performance in some settings.
Beyond Sentence-level Labels: Integrating Conversational Context and Personal Experience for Natural Emotional Expression (2026.findings-acl)

Copied to clipboard

Challenge: Existing systems rely on sentence-level labels, which fails to capture the subtle nuances of human affect.
Approach: They propose to use a large-scale, context-aware speech corpus derived from multi-speaker audiobooks to generate a speech that is human-like.
Outcome: The proposed model outperforms existing methods in terms of emotional expression accuracy and naturalness.
Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multimodal speech emotion recognition and sentiment analysis have not improved results due to their relatively simple fusion mechanisms and lack of proper cross-modal pretraining.
Approach: They propose a deep-fused audio-text bi-modal transformer with carefully designed cross-modal fusion mechanism and stage-wise cross-mod pretraining scheme to facilitate cross-modulation.
Outcome: The proposed method exceeds benchmarks on public IEMOCAP emotion and CMU-MOSEI sentiment datasets by a large margin.
EMTC: Multilabel Corpus in Movie Domain for Emotion Analysis in Conversational Text (L18-1)

Copied to clipboard

Challenge: Existing emotion corpora collected from twitters and use hashtags are limited in the number of characters.
Approach: They propose to build an emotion corpus based on conversational text data that includes 2.1 million utterances and is partly annotated by ourselves and independent annotators.
Outcome: The proposed corpus includes conversations from movies with more than 2.1 million utterances which are partly annotated by ourselves and independent annotators.
Emotion Impacts Speech Recognition Performance (N19-3)

Copied to clipboard

Challenge: Existing studies show that speech recognition systems depend on multiple factors including lexical content, speaker identity and dialect.
Approach: They propose a method that evaluates the impact of emotion on recognition even when manual transcripts are not available.
Outcome: The proposed method allows to evaluate the impact of emotion on recognition even when manual transcripts are not available.
Akan Cinematic Emotions (ACE): A Multimodal Multi-party Dataset for Emotion Recognition in Movie Dialogues (2025.findings-acl)

Copied to clipboard

Challenge: Akan Cinematic Emotions (AkaCE) is the first multimodal emotion dialogue dataset for an African language . it contains 385 emotion-labeled dialogues and 6162 utterances across audio, visual, and textual modalities, along with word-level prosodic prominence annotations.
Approach: They propose to use AkaCE to analyze African cinematic emotions using word-level prosodic prominence annotations.
Outcome: The Akan Cinematic Emotions (AkaCE) dataset addresses the significant lack of resources for low-resource languages in emotion recognition research.
MPDD: A Multi-Party Dialogue Dataset for Analysis of Emotions and Interpersonal Relationships (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets with emotion and relation labels for dialogues are limited.
Approach: They use a Chinese dialogue dataset to annotate emotions and interpersonal relationships on each utterance.
Outcome: The proposed dataset contains 25,548 utterances from 4,142 dialogues.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations