Challenge: Despite recent advances in speech-to-text translation, the impact of the emotion content has been overlooked.
Approach: They propose to use generative error correction (GER) to generate the translation based on the decoded N-best hypotheses and combine emotion and sentiment labels into the LLM finetuning process to enable the model to consider the emotion content.
Outcome: The proposed model can translate speech in English-Chinese using GER and emotion and sentiment labels.

Similar Papers

MELD-ST: An Emotion-aware Speech Translation Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Emotion plays a crucial role in human conversation.
Approach: They present a MELD-ST dataset for the emotion-aware speech translation task . they show that fine-tuning with emotion labels can enhance translation performance .
Outcome: The proposed dataset shows that fine tuning with emotion labels can improve translation performance in some settings.
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies have demonstrated that Large Language Models possess a form of emotional intelligence, capable of interpreting emotional stimuli in text.
Approach: They propose a method that translates speech characteristics into natural language descriptions and integrates them into LLMs to perform multimodal emotion analysis via text prompts.
Outcome: The proposed method outperforms baseline models that require structural modifications on two datasets showing significant improvements in emotion recognition accuracy.
Textless Speech Emotion Conversion using Discrete & Decomposed Representations (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for modifying emotion of speech are difficult because emotion affects all levels simultaneously.
Approach: They propose a method to convert a spoken language speech into a model of emotion . they use phonetic-content units, prosodic features, speaker, and emotion to modify the emotion a speech utterance has.
Outcome: The proposed method beats text-based systems in terms of perceived emotion and audio quality.
A Unified View on Emotion Representation in Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Recent studies show the presence of emotion concepts in the hidden state representations, but it’s unclear if the model has a robust representation consistent across different datasets.
Approach: They propose a unified view to understand emotion representation in Large Language Models by experimenting with diverse datasets and prompts.
Outcome: The proposed model can be interchanged between datasets with minimal impact on performance.
Adapting a Language Model for Controlled Affective Text Generation (2020.coling-main)

Copied to clipboard

Challenge: Existing models for affective text generation fail to capture emotional aspects of conversations without explicit affective information.
Approach: They propose to incorporate emotion as prior for the probabilistic state-of-the-art text generation model such as GPT-2 and incorporate emotion into the model to ensure grammatical correctness.
Outcome: The proposed model outperforms existing models in all intensities and is robust to human evaluations.
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling (2026.findings-acl)

Copied to clipboard

Challenge: Existing codecs optimize acoustic reconstruction, leaving emotion expressiveness insufficiently modeled at the representation level.
Approach: They propose an emotion-guided neural speech codec that preserves emotional information while maintaining semantic fidelity and prosodic naturalness.
Outcome: The proposed codec preserves emotional cues while maintaining semantic fidelity and prosodic naturalness.
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing speech-to-speech large language models rely on ASR transcription or use encoders to extract latent representations, weakening affective information and contextual coherence in multi-turn dialogues.
Approach: They propose a framework for speech-based empathetic response generation that captures turn-level affective states and dialogue-level emotional dynamics.
Outcome: The proposed framework outperforms baselines in automatic and human evaluations and remains robust across different Large Language Model (LLM) backbones.
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR).
Approach: They propose a multimodal LLM to receive source speech as extra input and reformat it as a cloze test with logits calibration to remove input information redundancy and simplify GER with clear instructions.
Outcome: The proposed model improves on 9 popular ASR datasets and is faster than vanilla GER.
On the Emotion Understanding of Synthesized Speech (2026.acl-long)

Copied to clipboard

Challenge: Existing models for emotion understanding do not capture fundamental features of synthesized speech.
Approach: They evaluate emotion recognition models on synthesized speech using SER models and generative models.
Outcome: The proposed model can't generalize to synthesized speech because of speech token prediction . generative models tend to infer emotion from textual semantics while ignoring paralinguistic cues.
Cross-Lingual Emotion Lexicon Induction using Representation Alignment in Low-Resource Settings (2020.coling-main)

Copied to clipboard

Challenge: Emotion lexicons provide information about associations between words and emotions.
Approach: They use crowdsourcing to annotate words with Plutchik's 8 basic emotions, providing binary labels.
Outcome: The proposed lexicons provide information about associations between words and emotions . the lexiconics are useful in emotional analyses of reviews, literary texts, and posts on social media .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations