IIIT-H TEMD Semi-Natural Emotional Speech Database from Professional Actors and Non-Actors (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing databases for emotion recognition are limited due to privacy and legal issues. |
| Approach: | They propose to collect emotional speech data from actors and non-actors using designed drama situations and annotate them manually using a hybrid strategy. |
| Outcome: | The proposed database is based on simulated parallel, semi-natural, and (near to) natural databases. |
Similar Papers
A Comparative Cross Language View On Acted Databases Portraying Basic Emotions Utilising Machine Learning (2022.lrec-1)
Copied to clipboard
| Challenge: | Since several decades emotional databases have been recorded by various laboratories. |
| Approach: | They propose to model similarity as performance in cross database machine learning experiments and to analyze a manually picked set of four acoustic features that represent different phonetic areas. |
| Outcome: | The proposed sets of features represent different phonetic areas and are comparable across languages. |
A Dataset for Speech Emotion Recognition in Greek Theatrical Plays (2022.lrec-1)
Copied to clipboard
| Challenge: | Speech Emotion Recognition (SER) is a task that is difficult to perform by humans due to subjectiveness of the emotional content. |
| Approach: | They propose to use GreThE to collect data for speech emotion recognition in Greek plays. |
| Outcome: | The proposed dataset contains utterances from various actors and plays, along with respective valence and arousal annotations. |
M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database (2022.acl-long)
Copied to clipboard
| Challenge: | Existing data resources to support multimodal affective analysis in dialogues are limited in scale and diversity. |
| Approach: | They propose a multimodal multi-scene multi-label Emotional Dialogue dataset, M3ED, which contains 990 dyadic emotional dialogues from 56 different TV series. |
| Outcome: | The proposed dataset contains 990 dyadic emotional dialogues from 56 different TV series, a total of 9,082 turns and 24,449 utterances. |
MELD-ST: An Emotion-aware Speech Translation Dataset (2024.findings-acl)
Copied to clipboard
Sirou Chen, Sakiko Yahata, Shuichiro Shimizu, Zhengdong Yang, Yihang Li, Chenhui Chu, Sadao Kurohashi
| Challenge: | Emotion plays a crucial role in human conversation. |
| Approach: | They present a MELD-ST dataset for the emotion-aware speech translation task . they show that fine-tuning with emotion labels can enhance translation performance . |
| Outcome: | The proposed dataset shows that fine tuning with emotion labels can improve translation performance in some settings. |
Beyond Sentence-level Labels: Integrating Conversational Context and Personal Experience for Natural Emotional Expression (2026.findings-acl)
Copied to clipboard
Haiyang Sun, Chenyang Le, Wei Wang, Leying Zhang, Chuang Li, Bing Han, Chenda Li, Mengxiao Bi, Yanmin Qian
| Challenge: | Existing systems rely on sentence-level labels, which fails to capture the subtle nuances of human affect. |
| Approach: | They propose to use a large-scale, context-aware speech corpus derived from multi-speaker audiobooks to generate a speech that is human-like. |
| Outcome: | The proposed model outperforms existing methods in terms of emotional expression accuracy and naturalness. |
Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to multimodal speech emotion recognition and sentiment analysis have not improved results due to their relatively simple fusion mechanisms and lack of proper cross-modal pretraining. |
| Approach: | They propose a deep-fused audio-text bi-modal transformer with carefully designed cross-modal fusion mechanism and stage-wise cross-mod pretraining scheme to facilitate cross-modulation. |
| Outcome: | The proposed method exceeds benchmarks on public IEMOCAP emotion and CMU-MOSEI sentiment datasets by a large margin. |
EMTC: Multilabel Corpus in Movie Domain for Emotion Analysis in Conversational Text (L18-1)
Copied to clipboard
| Challenge: | Existing emotion corpora collected from twitters and use hashtags are limited in the number of characters. |
| Approach: | They propose to build an emotion corpus based on conversational text data that includes 2.1 million utterances and is partly annotated by ourselves and independent annotators. |
| Outcome: | The proposed corpus includes conversations from movies with more than 2.1 million utterances which are partly annotated by ourselves and independent annotators. |
Emotion Impacts Speech Recognition Performance (N19-3)
Copied to clipboard
| Challenge: | Existing studies show that speech recognition systems depend on multiple factors including lexical content, speaker identity and dialect. |
| Approach: | They propose a method that evaluates the impact of emotion on recognition even when manual transcripts are not available. |
| Outcome: | The proposed method allows to evaluate the impact of emotion on recognition even when manual transcripts are not available. |
Akan Cinematic Emotions (ACE): A Multimodal Multi-party Dataset for Emotion Recognition in Movie Dialogues (2025.findings-acl)
Copied to clipboard
David Sasu, Zehui Wu, Ziwei Gong, Run Chen, Pengyuan Shi, Lin Ai, Julia Hirschberg, Natalie Schluter
| Challenge: | Akan Cinematic Emotions (AkaCE) is the first multimodal emotion dialogue dataset for an African language . it contains 385 emotion-labeled dialogues and 6162 utterances across audio, visual, and textual modalities, along with word-level prosodic prominence annotations. |
| Approach: | They propose to use AkaCE to analyze African cinematic emotions using word-level prosodic prominence annotations. |
| Outcome: | The Akan Cinematic Emotions (AkaCE) dataset addresses the significant lack of resources for low-resource languages in emotion recognition research. |
MPDD: A Multi-Party Dialogue Dataset for Analysis of Emotions and Interpersonal Relationships (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing datasets with emotion and relation labels for dialogues are limited. |
| Approach: | They use a Chinese dialogue dataset to annotate emotions and interpersonal relationships on each utterance. |
| Outcome: | The proposed dataset contains 25,548 utterances from 4,142 dialogues. |