Emotion Impacts Speech Recognition Performance (N19-3)

Copied to clipboard

Challenge: Existing studies show that speech recognition systems depend on multiple factors including lexical content, speaker identity and dialect.
Approach: They propose a method that evaluates the impact of emotion on recognition even when manual transcripts are not available.
Outcome: The proposed method allows to evaluate the impact of emotion on recognition even when manual transcripts are not available.

Similar Papers

The Correlation Between Emotion in Text and Speech Segments is Limited: A Cross-Modal Study (2026.findings-eacl)

Copied to clipboard

Challenge: a recent study has shown that text-to-speech systems can capture human-like emotion, but they lack the ability to predict emotion in speech.
Approach: They propose to use 8 large language models for identifying emotion in text and 2 audio models for emotion in speech to investigate the correlation between emotion and speech.
Outcome: The proposed models perform well on emotion recognition from situational text and audiobooks, but show weak correlation for Valence only.
Language-specific Effects on Automatic Speech Recognition Errors for World Englishes (2022.coling-1)

Copied to clipboard

Challenge: Existing systems are not able to meet the needs of speakers of different demographic groups.
Approach: They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors.
Outcome: The proposed system predicts certain errors from the phonological structure of a speaker’s native language.
Beyond Sentence-level Labels: Integrating Conversational Context and Personal Experience for Natural Emotional Expression (2026.findings-acl)

Copied to clipboard

Challenge: Existing systems rely on sentence-level labels, which fails to capture the subtle nuances of human affect.
Approach: They propose to use a large-scale, context-aware speech corpus derived from multi-speaker audiobooks to generate a speech that is human-like.
Outcome: The proposed model outperforms existing methods in terms of emotional expression accuracy and naturalness.
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance (2026.eacl-long)

Copied to clipboard

Challenge: LISTEN is a controlled benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding.
Approach: They propose a benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding.
Outcome: LISTEN shows that current LALMs largely "transcribe" rather than "listen" authors note that models underutilize acoustic cues while relying on lexical semantics .
Exploring the Role of Context in Utterance-level Emotion, Act and Intent Classification in Conversations: An Empirical Study (2021.findings-acl)

Copied to clipboard

Challenge: utterance-level dialogue understanding tasks are often performed at utterrance level and are often conjoined together under the umbrella of utterence-level dialog understanding.
Approach: They propose to use a contextual utterance-level dialogue understanding baseline as a strong framework for six dialogue-understanding tasks.
Outcome: The proposed framework can be easily adapted for other tasks for similar purposes.
On the Emotion Understanding of Synthesized Speech (2026.acl-long)

Copied to clipboard

Challenge: Existing models for emotion understanding do not capture fundamental features of synthesized speech.
Approach: They evaluate emotion recognition models on synthesized speech using SER models and generative models.
Outcome: The proposed model can't generalize to synthesized speech because of speech token prediction . generative models tend to infer emotion from textual semantics while ignoring paralinguistic cues.
A Comparative Cross Language View On Acted Databases Portraying Basic Emotions Utilising Machine Learning (2022.lrec-1)

Copied to clipboard

Challenge: Since several decades emotional databases have been recorded by various laboratories.
Approach: They propose to model similarity as performance in cross database machine learning experiments and to analyze a manually picked set of four acoustic features that represent different phonetic areas.
Outcome: The proposed sets of features represent different phonetic areas and are comparable across languages.
My Heart Skipped a Beat! Recognizing Expressions of Embodied Emotion in Natural Language (2024.naacl-long)

Copied to clipboard

Challenge: a new task is needed to recognize physical manifestations of emotions in natural language . physical manifestation of emotions affects not only our mental state but also our physical state .
Approach: They propose a task to recognize expressions of embodied emotion in natural language . they use body part mentions with human annotations to extract emotional manner expressions .
Outcome: The proposed model can train without gold data and improve performance with gold data.
LaERC-S: Improving LLM-based Emotion Recognition in Conversation with Speaker Characteristics (2025.coling-main)

Copied to clipboard

Challenge: Emotion recognition in conversation (ERC) is a task of discerning human emotions for each utterance within a conversation.
Approach: They propose a framework that uses large language models to analyze speaker characteristics . they use two-stage learning to make the models reason speaker characteristics and track emotion of the speaker .
Outcome: The proposed framework outperforms existing methods on three benchmark datasets.
Guilt by Association: Emotion Intensities in Lexical Representations (2021.emnlp-main)

Copied to clipboard

Challenge: linguistic models have a higher correlation with human ground truth ratings than labeled data . word vectors have often been evaluated on standard word relatedness benchmarks .
Approach: They propose to use unsupervised, supervised, and finally supervised methods to extract emotional associations from pretrained vectors and models.
Outcome: The proposed method shows higher correlation with ground truth ratings than state-of-the-art lexicons based on labeled data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations