| Challenge: | Emosical provides rich emotion annotations for musical films by inferring the background story of the characters. |
| Approach: | They propose to use a multimodal dataset of musical films to generate annotated emotion tags for each sample by inferring the background story of the characters. |
| Outcome: | The proposed dataset provides rich emotion annotations for musical films by inferring the background story of the characters. |
Similar Papers
A Dataset for Speech Emotion Recognition in Greek Theatrical Plays (2022.lrec-1)
Copied to clipboard
| Challenge: | Speech Emotion Recognition (SER) is a task that is difficult to perform by humans due to subjectiveness of the emotional content. |
| Approach: | They propose to use GreThE to collect data for speech emotion recognition in Greek plays. |
| Outcome: | The proposed dataset contains utterances from various actors and plays, along with respective valence and arousal annotations. |
Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS (2024.lrec-main)
Copied to clipboard
| Challenge: | We show that ascribing verbal descriptions to expressive audiovisual utterances is efficient and efficient. |
| Approach: | They propose a web app for ascribing verbal descriptions to expressive audiovisual utterances. |
| Outcome: | The proposed system can be deployed at a large scale to efficiently collect relevant verbal descriptions. |
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks fail to achieve ecological validity, signal clarity, and reliable fine-grained labeling in multimodal Emotion Recognition (MER) Existing datasets lack spontaneity of real-life interactions, resulting in poor quality and inconsistent data quality. |
| Approach: | They propose a bilingual benchmark to resolve limitations of ecological validity and noise in existing datasets by combining strictly filtered static slices with a dynamic Streaming Monologue subset. |
| Outcome: | EmoS provides trusted ground truth that captures continuous emotional evolution. |
Akan Cinematic Emotions (ACE): A Multimodal Multi-party Dataset for Emotion Recognition in Movie Dialogues (2025.findings-acl)
Copied to clipboard
David Sasu, Zehui Wu, Ziwei Gong, Run Chen, Pengyuan Shi, Lin Ai, Julia Hirschberg, Natalie Schluter
| Challenge: | Akan Cinematic Emotions (AkaCE) is the first multimodal emotion dialogue dataset for an African language . it contains 385 emotion-labeled dialogues and 6162 utterances across audio, visual, and textual modalities, along with word-level prosodic prominence annotations. |
| Approach: | They propose to use AkaCE to analyze African cinematic emotions using word-level prosodic prominence annotations. |
| Outcome: | The Akan Cinematic Emotions (AkaCE) dataset addresses the significant lack of resources for low-resource languages in emotion recognition research. |
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations (2026.findings-acl)
Copied to clipboard
Haoqin Sun, Jinghua Zhao, Xuechen Wang, Shiwan Zhao, Jiaming Zhou, Hui Wang, Xi Yang, Yequan Wang, Yonghua Lin
| Challenge: | Existing datasets face issues such as low quality, limited scale, and incomplete modalities, hindering model performance. |
| Approach: | They propose to use Chinese multimodal datasets to capture authentic emotional interplay from 19 professional actors. |
| Outcome: | The EmotionTalk dataset spans 23.6 hours of dyadic conversations across diverse scenarios. |
EMTC: Multilabel Corpus in Movie Domain for Emotion Analysis in Conversational Text (L18-1)
Copied to clipboard
| Challenge: | Existing emotion corpora collected from twitters and use hashtags are limited in the number of characters. |
| Approach: | They propose to build an emotion corpus based on conversational text data that includes 2.1 million utterances and is partly annotated by ourselves and independent annotators. |
| Outcome: | The proposed corpus includes conversations from movies with more than 2.1 million utterances which are partly annotated by ourselves and independent annotators. |
ELAL: An Emotion Lexicon for the Analysis of Alsatian Theatre Plays (2022.lrec-1)
Copied to clipboard
| Challenge: | a novel and manually corrected emotion lexicon is presented for Alsatian dialects . the dialects are used mainly orally and lack a stable and consensual spelling convention . |
| Approach: | They propose a novel and manually corrected emotion lexicon for Alsatian dialects . they use graphical variants of Alsalian lexical items to perform automatic emotion analysis . |
| Outcome: | The novel and manually corrected emotion lexicon is used to perform automatic emotion analysis in Alsatian theatre plays. |
Emo Pillars: Knowledge Distillation to Support Fine-Grained Context-Aware and Context-Less Emotion Classification (2025.findings-acl)
Copied to clipboard
| Challenge: | a recent study shows that sentiment analysis datasets lack context in which an opinion was expressed and are limited by a few emotion categories. |
| Approach: | They propose to ground an LLM-based model into a corpus of narratives to generate stories-character-centered utterances with unique contexts over 28 emotion classes. |
| Outcome: | The proposed model generates non-repetitive story-character-centered utterances with unique contexts over 28 emotion classes. |
EDA: Enriching Emotional Dialogue Acts using an Ensemble of Neural Annotators (2020.lrec-1)
Copied to clipboard
| Challenge: | Emotion recognition helps to build natural dialogue systems. |
| Approach: | They propose to use a recurrent neural model to annotate emotion corpora with dialogue act labels and an ensemble annotator to extract the final dialogue act label. |
| Outcome: | The proposed model annotates two accessible multi-modal emotion corpora with and without context and extracts the final dialogue act label. |
Folksonomication: Predicting Tags for Movies from Plot Synopses using Emotion Flow Encoded Neural Network (C18-1)
Copied to clipboard
| Challenge: | Existing systems that generate tags for movies can help users better retrieve movies based on their personal preferences and user profiles. |
| Approach: | They propose a neural network model that merges synopses and emotion flows to predict a set of movies' tags. |
| Outcome: | The proposed model outperforms a machine learning system by learning 18% more tags than the previous one. |