Challenge: Extensive research on spoken dialogue systems has advanced the development of intelligent voice assistants, but integration of role information within speech remains an underexplored area.
Approach: They propose a language-based spoken dialogue system that integrates role information within speech to generate contextually appropriate responses.
Outcome: The proposed architecture achieves speaker-specific responses, character understanding, and the generation of targeted replies in multi-party dialogue scenarios, surpassing existing spoken dialogue systems.

Similar Papers

A Persona-Aware LLM-Enhanced Framework for Multi-Session Personalized Dialogue Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing personalized dialogue models focus on dialogue history and personality information, reducing the responses’ consistency.
Approach: They propose a Persona-Aware LLM-enAnCEd(PALACE) framework that generates responses consistent with dialogue history and personality information across multiple sessions to engage users’ interest in the dialogue.
Outcome: The proposed framework outperforms the state-of-the-art methods in automatic and human evaluation metrics on the MSC and DuLeMon datasets.
Post Persona Alignment for Multi-Session Dialogue Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for multi-session persona-based dialogue generation typically retrieve persona information before response generation, which can constrain diversity and result in generic outputs.
Approach: They propose a two-stage framework that reverses the process of retrieving persona information before response generation.
Outcome: Experiments on multi-session persona-based dialogue data show that the proposed framework outperforms existing methods in consistency, diversity, and persona relevance.
Learning to Improve Persona Consistency in Multi-party Dialogue Generation via Text Knowledge Enhancement (2022.coling-1)

Copied to clipboard

Challenge: Existing methods suffer from incomprehensive persona tags that have unique and obscure meanings to describe human’s personality.
Approach: They propose a graph convolution network model with addressee selecting mechanism that integrates personas, dialogue utterances, and external text knowledge in a unified graph.
Outcome: The proposed model outperforms baselines by large margins and improves persona consistency in the generated responses.
Pre-training Multi-party Dialogue Models with Latent Discourse Inference (2023.acl-long)

Copied to clipboard

Challenge: Existing studies have failed to scale up the pre-training process by putting aside unlabeled data . et al., 2019: multi-party dialogues are more difficult for models to understand since they involve multiple interlocutors resulting in interweaving reply-to relations and information flows.
Approach: They propose to treat discourse structures as latent variables and jointly infer them to pre-train a model that understands the discourse structure of multi-party dialogues.
Outcome: The proposed model outperforms baselines and achieves state-of-the-art results on multiple downstream tasks.
Speaker-Aware Discourse Parsing on Multi-Party Dialogues (2022.coling-1)

Copied to clipboard

Challenge: Discourse parsing on multi-party dialogues is an important but difficult task in dialogue systems and conversational analysis.
Approach: They propose a speaker-aware model for parsing on multi-party dialogues using interaction features between different speakers.
Outcome: The proposed model achieves the best-reported performance on two standard benchmark datasets.
Towards Neural Speaker Modeling in Multi-Party Conversation: The Task, Dataset, and Models (L18-1)

Copied to clipboard

Challenge: Existing methods for speaker modeling are based on hand-crafted statistics and ad hoc to a certain application.
Approach: They propose to use speaker classification as a surrogate task for general speaker modeling and collect massive data to facilitate research in this direction.
Outcome: The proposed models outperform the existing models and are feasible with speaker identity information.
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue (2025.naacl-long)

Copied to clipboard

Challenge: Existing dialogue systems focus on brief single-session interactions, neglecting real-world needs for long-term companionship and personalized interactions.
Approach: They propose a model-agnostic framework for long-term dialogue agents . they use event summary and persona management to enable reasoning .
Outcome: The proposed framework incorporates three independently tunable modules dedicated to event perception, persona extraction, and response generation.
I know you are different! Towards Persona Driven Knowledge-infused Dialogue Assistant (2026.eacl-long)

Copied to clipboard

Challenge: Task-Oriented Dialogue (TOD) systems often fall short in delivering personalized, context-rich responses, especially in low-resource, code-mixed, and multimodal settings like Hinglish.
Approach: They propose a Hinglish multimodal, multidomain, persona-based TOD dataset that captures user-agent interactions across text and visual modalities.
Outcome: The proposed framework outperforms standard and ablated models in Hinglish and Hinglanish.
Who is Speaking? Speaker-Aware Multiparty Dialogue Act Classification (2023.findings-emnlp)

Copied to clipboard

Challenge: Identifying how speakers interact with each other in a conversation is difficult when more than two interlocutors take part in . To overcome this challenge, we propose to explicitly add speaker awareness to each utterance representation.
Approach: They propose to add speaker awareness to each utterance representation to model how each speaker is behaving within the local context of a conversation.
Outcome: The proposed approach is able to model multiparticipant and dyadic conversations on the MRDA and SwDA datasets and shows that it is more efficient than previous approaches.
Is ChatGPT a Good Multi-Party Conversation Solver? (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are powerful tools for multi-party conversations, but their capacity to handle multi-parties remains unexplored.
Approach: They propose to evaluate ChatGPT and GPT-4's zero-shot learning capabilities within the context of multi-party conversations (MPCs) they also propose to incorporate MPC structures, encompassing both speaker and addressee architecture.
Outcome: The proposed models perform poorly on a number of MPC tasks while GPT-4 performs well on speaker and addressee architecture.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations