Sirou Chen, Sakiko Yahata, Shuichiro Shimizu, Zhengdong Yang, Yihang Li, Chenhui Chu, Sadao Kurohashi
| Challenge: | Emotion plays a crucial role in human conversation. |
| Approach: | They present a MELD-ST dataset for the emotion-aware speech translation task . they show that fine-tuning with emotion labels can enhance translation performance . |
| Outcome: | The proposed dataset shows that fine tuning with emotion labels can improve translation performance in some settings. |
Similar Papers
Generative Error Correction for Emotion-aware Speech-to-text Translation (2025.findings-acl)
Copied to clipboard
| Challenge: | Despite recent advances in speech-to-text translation, the impact of the emotion content has been overlooked. |
| Approach: | They propose to use generative error correction (GER) to generate the translation based on the decoded N-best hypotheses and combine emotion and sentiment labels into the LLM finetuning process to enable the model to consider the emotion content. |
| Outcome: | The proposed model can translate speech in English-Chinese using GER and emotion and sentiment labels. |
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations (P19-1)
Copied to clipboard
| Challenge: | Emotion recognition in conversations has gained popularity due to its potential applications. Until now, a large multimodal multi-party emotional conversational database containing more than two speakers per dialogue was missing. |
| Approach: | They propose to extend and enhance EmotionLines by combining 13,000 utterances from Friends dialogues with emotion and sentiment labels. |
| Outcome: | The proposed dataset contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends. |
Sentiment Analysis for Emotional Speech Synthesis in a News Dialogue System (2020.coling-main)
Copied to clipboard
| Challenge: | In smart speakers and conversational robots, the demand for expressive speech synthesis has increased. |
| Approach: | They propose to annotate a news dataset with emotion labels for each sentence and to evaluate its effectiveness using the constructed dataset. |
| Outcome: | The proposed method improves the performance of the proposed model by preferentially annotating news articles with low confidence in the human-in-the-loop machine learning framework. |
An Emotional Mess! Deciding on a Framework for Building a Dutch Emotion-Annotated Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing frameworks for emotion recognition are limited and do not allow for categorical versus dimensional oppositions. |
| Approach: | They propose to use the emotions joy, love, anger, sadness and fear as well as dimensional models to annotate texts from different domains and topics. |
| Outcome: | The proposed frameworks are well-suited to annotate texts from different domains and topics, but the connotation of the labels strongly depends on the origin of the texts. |
Textless Speech Emotion Conversion using Discrete & Decomposed Representations (2022.emnlp-main)
Copied to clipboard
Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu Anh Nguyen, Morgan Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, Yossi Adi
| Challenge: | Existing methods for modifying emotion of speech are difficult because emotion affects all levels simultaneously. |
| Approach: | They propose a method to convert a spoken language speech into a model of emotion . they use phonetic-content units, prosodic features, speaker, and emotion to modify the emotion a speech utterance has. |
| Outcome: | The proposed method beats text-based systems in terms of perceived emotion and audio quality. |
On the Emotion Understanding of Synthesized Speech (2026.acl-long)
Copied to clipboard
Yuan Ge, Haishu Zhao, AoKai Hao, Junxiang Zhang, Bei Li, Xiaoqian Liu, Chenglong Wang, Jianjin Wang, Bingsen Zhou, Bingyu Liu, JingBo Zhu, Zhengtao Yu, Tong Xiao
| Challenge: | Existing models for emotion understanding do not capture fundamental features of synthesized speech. |
| Approach: | They evaluate emotion recognition models on synthesized speech using SER models and generative models. |
| Outcome: | The proposed model can't generalize to synthesized speech because of speech token prediction . generative models tend to infer emotion from textual semantics while ignoring paralinguistic cues. |
EmotionLines: An Emotion Corpus of Multi-Party Conversations (L18-1)
Copied to clipboard
| Challenge: | Emotion is a critical characteristic to distinguish people from machines. |
| Approach: | They propose a dataset with emotions labeling on all utterances in each dialogue . they use Friends TV scripts and Facebook messenger dialogues to collect the data . |
| Outcome: | The proposed dataset is the first with emotions labeling on all utterances in each dialogue based on their textual content. |
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations (2026.findings-acl)
Copied to clipboard
Haoqin Sun, Jinghua Zhao, Xuechen Wang, Shiwan Zhao, Jiaming Zhou, Hui Wang, Xi Yang, Yequan Wang, Yonghua Lin
| Challenge: | Existing datasets face issues such as low quality, limited scale, and incomplete modalities, hindering model performance. |
| Approach: | They propose to use Chinese multimodal datasets to capture authentic emotional interplay from 19 professional actors. |
| Outcome: | The EmotionTalk dataset spans 23.6 hours of dyadic conversations across diverse scenarios. |
Learning and Evaluating Emotion Lexicons for 91 Languages (2020.acl-main)
Copied to clipboard
| Challenge: | Emotion lexicons describe the affective meaning of words but are limited in coverage for most languages. |
| Approach: | They propose a method for creating arbitrarily large emotion lexicons for any target language. |
| Outcome: | The proposed method exceeds human reliability for some languages and variables. |
MEISD: A Multimodal Multi-Label Emotion, Intensity and Sentiment Dialogue Dataset for Emotion Recognition and Sentiment Analysis in Conversations (2020.coling-main)
Copied to clipboard
| Challenge: | Emotion and sentiment classification in dialogues has gained popularity in recent times . a number of datasets are imbalanced in representing different emotions and consist of an only single emotion. |
| Approach: | They propose to use a dataset to analyze emotions and sentiments in dialogues . they use text, audio and video to identify the correct emotions with the appropriate intensity and sentiment in an utterance of a dialogue . |
| Outcome: | The proposed datasets are balanced in representing different emotions and consist of only one emotion. |