Challenge: Recent advances in multi-turn voice interaction models have improved user-model communication, but whether open-source models share this ability remains unexplored.
Approach: They propose to use ContextDialog to evaluate open-source interaction models' ability to recall past utterances to identify key limitations.
Outcome: The proposed model retains and recalls past utterances better than closed-source models, but still struggles with questions about past . findings highlight key limitations in open-source model and suggest ways to improve memory retention and retrieval robustness.

Similar Papers

Beyond Goldfish Memory: Long-Term Open-Domain Conversation (2022.acl-long)

Copied to clipboard

Challenge: Despite recent improvements in open-domain dialogue models, state of the art models are trained and evaluated on short conversations with little context.
Approach: They propose to use retrieval-augmented methods to summarize and recall past conversations to improve their models.
Outcome: The proposed models outperform the current state-of-the-art models on human-human chat sessions in both automatic and human evaluations.
Do Neural Dialog Systems Use the Conversation History Effectively? An Empirical Study (P19-1)

Copied to clipboard

Challenge: Neural generative models are becoming more popular when building conversational agents.
Approach: They propose to study the sensitivity of neural dialog models to unnatural perturbations . they experiment with 10 different types of perturbations on 4 multi-turn dialog datasets .
Outcome: The proposed model is sensitive to unnatural changes or perturbations on 4 multi-turn dialog datasets.
Recent Advances in Speech Language Models: A Survey (2025.acl-long)

Copied to clipboard

Challenge: Text-based Large Language Models (LLMs) are a promising solution for end-to-end speech interaction.
Approach: They propose to build a framework that allows users to input text and translate it into speech . they propose to use a text-only LLM and a "textto-speech" framework to generate a response based on this transcription .
Outcome: The survey offers an overview of recent approaches to building SpeechLMs . it outlines core architectural components, training methodologies, evaluation strategies and challenges .
Evaluating Open-Source ASR Systems: Performance Across Diverse Audio Conditions and Error Correction Methods (2025.coling-main)

Copied to clipboard

Challenge: Automated speech recognition (ASR) systems are able to transcribe spontaneous human conversations with high accuracy.
Approach: They evaluate the accuracy of open source automatic speech recognition systems across conversational speech datasets and explore the potential of ASR ensembling and post-ASR correction methods to improve transcription accuracy.
Outcome: The proposed methods highlight the need for robust error correction techniques and address demographic biases to enhance ASR performance and inclusivity.
How to Represent Context Better? An Empirical Study on Context Modeling for Multi-turn Response Selection (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing work on building a conversational system for open domain human-machine conversation is attracting more attention . early models concatenate all utterances or independently encode each dialogue turn, which may lead to an inadequate understanding of dialogue status.
Approach: They propose to use a turn-aware context modeling layer to adapt existing models . they propose to model multi-turn contexts from the perspective of sequential relationship, local relationship, and query-alike manner .
Outcome: The proposed method can be adapted to several advanced response selection models.
Session-level Language Modeling for Conversational Speech (D18-1)

Copied to clipboard

Challenge: Xiong et al., 2017) generalizes language models for conversational speech recognition . recurrent neural networks (RNNs) read a list of words sequentially and predict the next word at each position.
Approach: They propose to generalize language models for conversational speech recognition to capture conversation-level phenomena such as adjacency pairs, lexical entrainment, and topical coherence.
Outcome: The proposed model reduces perplexity and improves word error rate over standard models in the conversational telephone speech domain.
A Dynamic Speaker Model for Conversational Interactions (N19-1)

Copied to clipboard

Challenge: a neural model for characterizing individual differences in speakers is shown to be useful in human-computer interaction and dialog act prediction.
Approach: They propose a neural model for learning a dynamically updated speaker embedding in a conversational context.
Outcome: The proposed model is used for content ranking and dialog act prediction in human-human conversations.
Evaluating the Long-Term Memory of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have not thoroughly investigated the memory performance of large language models in long-term tasks.
Approach: They propose a dataset to evaluate the long-term memory capabilities of large language models.
Outcome: The proposed model exhibits memory preferences across different categories of information.
Keep Me Updated! Memory Management in Long-term Conversations (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies do not deal with cases where memorized information is outdated, which may cause confusion in later conversations.
Approach: They propose a task where bots keep track of and bring up the latest information about users while conversing through multiple sessions.
Outcome: The proposed method outperforms baselines that leave the stored memory unchanged in terms of engagingness and humanness, and a larger performance gap in the later sessions.
InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model (2025.findings-emnlp)

Copied to clipboard

Challenge: Spoken Dialogue models face challenges in handling nuanced interactional phenomena, such as interruptions and backchannels.
Approach: They propose to use a 150-hour English speech interaction dialogue dataset to empower spoken dialogue models with nuanced real-time interaction capabilities.
Outcome: The proposed dataset trains and evaluates a speech understanding model that classifies key interactional events directly from audio.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations