| Challenge: | Xiong et al., 2017) generalizes language models for conversational speech recognition . recurrent neural networks (RNNs) read a list of words sequentially and predict the next word at each position. |
| Approach: | They propose to generalize language models for conversational speech recognition to capture conversation-level phenomena such as adjacency pairs, lexical entrainment, and topical coherence. |
| Outcome: | The proposed model reduces perplexity and improves word error rate over standard models in the conversational telephone speech domain. |
Similar Papers
Acoustic-to-Word Models with Conversational Context Information (N19-1)
Copied to clipboard
| Challenge: | Existing speech recognition models are built at a sentence level, and therefore it may not capture conversational context information. |
| Approach: | They propose a direct acoustic-to-word, end-to end speech recognition model that integrates a conversational context with other available information and directly recognizes words from speech. |
| Outcome: | The proposed model outperforms a standard end-to-end speech recognition system on the Switchboard conversational speech corpus and shows that it is more accurate than existing models. |
Beyond Goldfish Memory: Long-Term Open-Domain Conversation (2022.acl-long)
Copied to clipboard
| Challenge: | Despite recent improvements in open-domain dialogue models, state of the art models are trained and evaluated on short conversations with little context. |
| Approach: | They propose to use retrieval-augmented methods to summarize and recall past conversations to improve their models. |
| Outcome: | The proposed models outperform the current state-of-the-art models on human-human chat sessions in both automatic and human evaluations. |
A Dynamic Speaker Model for Conversational Interactions (N19-1)
Copied to clipboard
| Challenge: | a neural model for characterizing individual differences in speakers is shown to be useful in human-computer interaction and dialog act prediction. |
| Approach: | They propose a neural model for learning a dynamically updated speaker embedding in a conversational context. |
| Outcome: | The proposed model is used for content ranking and dialog act prediction in human-human conversations. |
Spoken Conversational Agents with Large Language Models (2025.emnlp-tutorials)
Copied to clipboard
| Challenge: | This tutorial focuses on the evolution of voice-native LLMs . it reviews the adaptation of text LLM to audio, cross-modal alignment, and joint speech–text training . |
| Approach: | This tutorial examines the evolution of voice-native LLMs in conversational agents . it compares cascaded and voice-based LLM systems to end-to-end retrieval-and vision-grounded systems . |
| Outcome: | This tutorial examines the evolution of voice-native LLMs . it compares the performance of voice assistants to current open-domain agents . |
Adaptation of Hierarchical Structured Models for Speech Act Recognition in Asynchronous Conversation (N19-1)
Copied to clipboard
| Challenge: | asynchronous domains lack large labeled datasets to train an effective speech act recognition model. |
| Approach: | They propose methods to leverage abundant unlabeled conversational data and available labeled data from synchronous domains to train an effective SAR model. |
| Outcome: | The proposed method outperforms existing methods when trained on in-domain data only. |
Chameleon: A Language Model Adaptation Toolkit for Automatic Speech Recognition of Conversational Speech (D19-3)
Copied to clipboard
| Challenge: | Language model adaptation (LMA) is a promising solution for conversational speech recognition systems. |
| Approach: | They propose to use language model adaptation techniques to adapt language models to conversational speech recognition. |
| Outcome: | The proposed toolkit compares state-of-the-art language model adaptation techniques in conversational speech recognition tasks. |
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in multi-turn voice interaction models have improved user-model communication, but whether open-source models share this ability remains unexplored. |
| Approach: | They propose to use ContextDialog to evaluate open-source interaction models' ability to recall past utterances to identify key limitations. |
| Outcome: | The proposed model retains and recalls past utterances better than closed-source models, but still struggles with questions about past . findings highlight key limitations in open-source model and suggest ways to improve memory retention and retrieval robustness. |
Utterance-level Detection Framework for LLM-Involved Content Detection in Conversational Setting (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing methods focus on static, document-level content, overlooking the dynamic nature of dialogues. |
| Approach: | They propose an utterance-level detection framework which integrates features from individual and combined analysis of dialogue participants’ responses to detect LLM-generated text under conversational setting. |
| Outcome: | The proposed framework achieves 98.14% accuracy with high inference speed and extensive results on different models and settings. |
Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding (2024.emnlp-main)
Copied to clipboard
| Challenge: | despite its importance, there has been limited research on conversational grounding in recent years . pre-trained language models have been costly and time-consuming to evaluate . |
| Approach: | They evaluate the performance of large language models in various aspects of conversational grounding . they propose ways to enhance the capabilities of the models that lag in this aspect . |
| Outcome: | The proposed model performance is based on pre-trained language models and a large pre-training dataset. |
Conversational Memory Network for Emotion Recognition in Dyadic Dialogue Videos (N18-1)
Copied to clipboard
Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria, Louis-Philippe Morency, Roger Zimmermann
| Challenge: | Existing methods for recognizing emotions in conversations ignore inter-speaker dependency relations . dyadic conversations are a form of dialogue between two entities . |
| Approach: | They propose a deep neural framework which leverages contextual information from the conversation history to model past utterances of each speaker into memories. |
| Outcome: | The proposed framework improves by 3 4% over the state-of-the-art in recognizing emotions in dyadic conversational videos. |