Session-level Language Modeling for Conversational Speech (D18-1)

Copied to clipboard

Challenge: Xiong et al., 2017) generalizes language models for conversational speech recognition . recurrent neural networks (RNNs) read a list of words sequentially and predict the next word at each position.
Approach: They propose to generalize language models for conversational speech recognition to capture conversation-level phenomena such as adjacency pairs, lexical entrainment, and topical coherence.
Outcome: The proposed model reduces perplexity and improves word error rate over standard models in the conversational telephone speech domain.

Similar Papers

Acoustic-to-Word Models with Conversational Context Information (N19-1)

Copied to clipboard

Challenge: Existing speech recognition models are built at a sentence level, and therefore it may not capture conversational context information.
Approach: They propose a direct acoustic-to-word, end-to end speech recognition model that integrates a conversational context with other available information and directly recognizes words from speech.
Outcome: The proposed model outperforms a standard end-to-end speech recognition system on the Switchboard conversational speech corpus and shows that it is more accurate than existing models.
Beyond Goldfish Memory: Long-Term Open-Domain Conversation (2022.acl-long)

Copied to clipboard

Challenge: Despite recent improvements in open-domain dialogue models, state of the art models are trained and evaluated on short conversations with little context.
Approach: They propose to use retrieval-augmented methods to summarize and recall past conversations to improve their models.
Outcome: The proposed models outperform the current state-of-the-art models on human-human chat sessions in both automatic and human evaluations.
A Dynamic Speaker Model for Conversational Interactions (N19-1)

Copied to clipboard

Challenge: a neural model for characterizing individual differences in speakers is shown to be useful in human-computer interaction and dialog act prediction.
Approach: They propose a neural model for learning a dynamically updated speaker embedding in a conversational context.
Outcome: The proposed model is used for content ranking and dialog act prediction in human-human conversations.
Spoken Conversational Agents with Large Language Models (2025.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial focuses on the evolution of voice-native LLMs . it reviews the adaptation of text LLM to audio, cross-modal alignment, and joint speech–text training .
Approach: This tutorial examines the evolution of voice-native LLMs in conversational agents . it compares cascaded and voice-based LLM systems to end-to-end retrieval-and vision-grounded systems .
Outcome: This tutorial examines the evolution of voice-native LLMs . it compares the performance of voice assistants to current open-domain agents .
Adaptation of Hierarchical Structured Models for Speech Act Recognition in Asynchronous Conversation (N19-1)

Copied to clipboard

Challenge: asynchronous domains lack large labeled datasets to train an effective speech act recognition model.
Approach: They propose methods to leverage abundant unlabeled conversational data and available labeled data from synchronous domains to train an effective SAR model.
Outcome: The proposed method outperforms existing methods when trained on in-domain data only.
Chameleon: A Language Model Adaptation Toolkit for Automatic Speech Recognition of Conversational Speech (D19-3)

Copied to clipboard

Challenge: Language model adaptation (LMA) is a promising solution for conversational speech recognition systems.
Approach: They propose to use language model adaptation techniques to adapt language models to conversational speech recognition.
Outcome: The proposed toolkit compares state-of-the-art language model adaptation techniques in conversational speech recognition tasks.
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in multi-turn voice interaction models have improved user-model communication, but whether open-source models share this ability remains unexplored.
Approach: They propose to use ContextDialog to evaluate open-source interaction models' ability to recall past utterances to identify key limitations.
Outcome: The proposed model retains and recalls past utterances better than closed-source models, but still struggles with questions about past . findings highlight key limitations in open-source model and suggest ways to improve memory retention and retrieval robustness.
Utterance-level Detection Framework for LLM-Involved Content Detection in Conversational Setting (2026.eacl-long)

Copied to clipboard

Challenge: Existing methods focus on static, document-level content, overlooking the dynamic nature of dialogues.
Approach: They propose an utterance-level detection framework which integrates features from individual and combined analysis of dialogue participants’ responses to detect LLM-generated text under conversational setting.
Outcome: The proposed framework achieves 98.14% accuracy with high inference speed and extensive results on different models and settings.
Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding (2024.emnlp-main)

Copied to clipboard

Challenge: despite its importance, there has been limited research on conversational grounding in recent years . pre-trained language models have been costly and time-consuming to evaluate .
Approach: They evaluate the performance of large language models in various aspects of conversational grounding . they propose ways to enhance the capabilities of the models that lag in this aspect .
Outcome: The proposed model performance is based on pre-trained language models and a large pre-training dataset.
Conversational Memory Network for Emotion Recognition in Dyadic Dialogue Videos (N18-1)

Copied to clipboard

Challenge: Existing methods for recognizing emotions in conversations ignore inter-speaker dependency relations . dyadic conversations are a form of dialogue between two entities .
Approach: They propose a deep neural framework which leverages contextual information from the conversation history to model past utterances of each speaker into memories.
Outcome: The proposed framework improves by 3 4% over the state-of-the-art in recognizing emotions in dyadic conversational videos.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations