| Challenge: | Existing studies show that speech recognition systems depend on multiple factors including lexical content, speaker identity and dialect. |
| Approach: | They propose a method that evaluates the impact of emotion on recognition even when manual transcripts are not available. |
| Outcome: | The proposed method allows to evaluate the impact of emotion on recognition even when manual transcripts are not available. |
Similar Papers
The Correlation Between Emotion in Text and Speech Segments is Limited: A Cross-Modal Study (2026.findings-eacl)
Copied to clipboard
| Challenge: | a recent study has shown that text-to-speech systems can capture human-like emotion, but they lack the ability to predict emotion in speech. |
| Approach: | They propose to use 8 large language models for identifying emotion in text and 2 audio models for emotion in speech to investigate the correlation between emotion and speech. |
| Outcome: | The proposed models perform well on emotion recognition from situational text and audiobooks, but show weak correlation for Valence only. |
Language-specific Effects on Automatic Speech Recognition Errors for World Englishes (2022.coling-1)
Copied to clipboard
| Challenge: | Existing systems are not able to meet the needs of speakers of different demographic groups. |
| Approach: | They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors. |
| Outcome: | The proposed system predicts certain errors from the phonological structure of a speaker’s native language. |
Beyond Sentence-level Labels: Integrating Conversational Context and Personal Experience for Natural Emotional Expression (2026.findings-acl)
Copied to clipboard
Haiyang Sun, Chenyang Le, Wei Wang, Leying Zhang, Chuang Li, Bing Han, Chenda Li, Mengxiao Bi, Yanmin Qian
| Challenge: | Existing systems rely on sentence-level labels, which fails to capture the subtle nuances of human affect. |
| Approach: | They propose to use a large-scale, context-aware speech corpus derived from multi-speaker audiobooks to generate a speech that is human-like. |
| Outcome: | The proposed model outperforms existing methods in terms of emotional expression accuracy and naturalness. |
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance (2026.eacl-long)
Copied to clipboard
| Challenge: | LISTEN is a controlled benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding. |
| Approach: | They propose a benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding. |
| Outcome: | LISTEN shows that current LALMs largely "transcribe" rather than "listen" authors note that models underutilize acoustic cues while relying on lexical semantics . |
Exploring the Role of Context in Utterance-level Emotion, Act and Intent Classification in Conversations: An Empirical Study (2021.findings-acl)
Copied to clipboard
| Challenge: | utterance-level dialogue understanding tasks are often performed at utterrance level and are often conjoined together under the umbrella of utterence-level dialog understanding. |
| Approach: | They propose to use a contextual utterance-level dialogue understanding baseline as a strong framework for six dialogue-understanding tasks. |
| Outcome: | The proposed framework can be easily adapted for other tasks for similar purposes. |
On the Emotion Understanding of Synthesized Speech (2026.acl-long)
Copied to clipboard
Yuan Ge, Haishu Zhao, AoKai Hao, Junxiang Zhang, Bei Li, Xiaoqian Liu, Chenglong Wang, Jianjin Wang, Bingsen Zhou, Bingyu Liu, JingBo Zhu, Zhengtao Yu, Tong Xiao
| Challenge: | Existing models for emotion understanding do not capture fundamental features of synthesized speech. |
| Approach: | They evaluate emotion recognition models on synthesized speech using SER models and generative models. |
| Outcome: | The proposed model can't generalize to synthesized speech because of speech token prediction . generative models tend to infer emotion from textual semantics while ignoring paralinguistic cues. |
A Comparative Cross Language View On Acted Databases Portraying Basic Emotions Utilising Machine Learning (2022.lrec-1)
Copied to clipboard
| Challenge: | Since several decades emotional databases have been recorded by various laboratories. |
| Approach: | They propose to model similarity as performance in cross database machine learning experiments and to analyze a manually picked set of four acoustic features that represent different phonetic areas. |
| Outcome: | The proposed sets of features represent different phonetic areas and are comparable across languages. |
My Heart Skipped a Beat! Recognizing Expressions of Embodied Emotion in Natural Language (2024.naacl-long)
Copied to clipboard
| Challenge: | a new task is needed to recognize physical manifestations of emotions in natural language . physical manifestation of emotions affects not only our mental state but also our physical state . |
| Approach: | They propose a task to recognize expressions of embodied emotion in natural language . they use body part mentions with human annotations to extract emotional manner expressions . |
| Outcome: | The proposed model can train without gold data and improve performance with gold data. |
LaERC-S: Improving LLM-based Emotion Recognition in Conversation with Speaker Characteristics (2025.coling-main)
Copied to clipboard
| Challenge: | Emotion recognition in conversation (ERC) is a task of discerning human emotions for each utterance within a conversation. |
| Approach: | They propose a framework that uses large language models to analyze speaker characteristics . they use two-stage learning to make the models reason speaker characteristics and track emotion of the speaker . |
| Outcome: | The proposed framework outperforms existing methods on three benchmark datasets. |
Guilt by Association: Emotion Intensities in Lexical Representations (2021.emnlp-main)
Copied to clipboard
| Challenge: | linguistic models have a higher correlation with human ground truth ratings than labeled data . word vectors have often been evaluated on standard word relatedness benchmarks . |
| Approach: | They propose to use unsupervised, supervised, and finally supervised methods to extract emotional associations from pretrained vectors and models. |
| Outcome: | The proposed method shows higher correlation with ground truth ratings than state-of-the-art lexicons based on labeled data. |