Training Millions of Personalized Dialogue Agents (D18-1)

Copied to clipboard

Challenge: Current dialogue systems fail at being engaging for users when trained end-to-end without relying on proactive reengaging scripted strategies.
Approach: They propose a dataset that provides 5 million personas and 700 million person-based dialogues.
Outcome: The proposed dataset provides 5 million personas and 700 million person-based dialogues.

Similar Papers

Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations (2024.lrec-main)

Copied to clipboard

Challenge: Personalization is a multifaceted process that requires multiple definitions and varies between individuals.
Approach: They propose to systemically survey the recent landscape of personalized dialogue generation including the datasets employed, methodologies developed, and evaluation metrics applied.
Outcome: The proposed model can generate fluent and coherent responses to human queries in a language-based conversational agent.
“In-Dialogues We Learn”: Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to personalized dialogue generate pre-defined profiles that are time-consuming and labor-intensive to create.
Approach: They propose a framework that leverages dialogue history to characterize personas without pre-defined profiles.
Outcome: The proposed framework improves BLEU and ROUGE scores on three datasets and human evaluations further validate the proposed method.
Beyond Candidates : Adaptive Dialogue Agent Utilizing Persona and Knowledge (2023.findings-emnlp)

Copied to clipboard

Challenge: a previous study suggested that human dialogue systems ground persona and knowledge but they require incomplete candidate sets.
Approach: They propose an adaptive dialogue agent that uses persona and knowledge without candidate sets . their model generates consistent and relevant persona descriptions and identifies relevant knowledge .
Outcome: The proposed model outperforms baselines that ground persona and knowledge candidates even with fragmentary information.
Personalizing Dialogue Agents via Meta-Learning (P19-1)

Copied to clipboard

Challenge: Existing personalized dialogue models use human designed persona descriptions to improve dialogue consistency.
Approach: They propose to extend Model-Agnostic Meta-Learning (MAML) to personalized dialogue learning without using persona descriptions.
Outcome: The proposed model outperforms baseline models in terms of human-evaluated fluency and consistency on a persona-chat dataset.
A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation (2022.acl-long)

Copied to clipboard

Challenge: Existing models for introducing explicit personas are expensive due to their expensive collection costs.
Approach: They propose a data manipulation method which is model-agnostic to be packed with any persona-based dialogue generation model to improve their performance.
Outcome: The proposed method is model-agnostic to be packed with any persona-based dialogue generation model to improve their performance.
Data Collection and End-to-End Learning for Conversational AI (D19-2)

Copied to clipboard

Challenge: tutorial aims to familiarise research community with recent advances in statistical dialogue systems . focus of tutorial is on learning end-to-end from data and their relation to more common modular systems.
Approach: This tutorial aims to familiarise the research community with the latest advances in statistical dialogue systems . the focus of the tutorial is on recently introduced end-to-end learning for dialogue systems and their relation to more common modular systems.
Outcome: This tutorial aims to familiarise the research community with the recent advances in statistical dialogue systems for open-domain and task-based dialogue paradigms.
Learning to Predict Persona Information for Dialogue Personalization without Explicit Persona Description (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to personalize dialogue agents rely on explicit persona descriptions during inference, which severely limits their application in real-world scenarios.
Approach: They propose a method that learns to predict persona information based on the dialogue history to personalize dialogue agents without relying on explicit persona descriptions during inference.
Outcome: The proposed method improves the consistency and engagingness of generated responses when conditioning on the predicted profile of the dialogue agent.
LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming (2023.acl-long)

Copied to clipboard

Challenge: a recent study shows that open-domain dialogue systems are not able to perform well in fast-growing scenarios such as live streaming due to the domain gap between online-post constructed data and those required in downstream conversational tasks.
Approach: They propose to train a conversational agent based on large social media datasets with multiple domains to improve response in live streaming scenarios.
Outcome: The proposed model improves response modeling and addressee recognition in live open-domain scenarios.
Large Scale Multi-Actor Generative Dialog Modeling (2020.acl-main)

Copied to clipboard

Challenge: Non-goal oriented dialog agents typically exhibit inconsistent personality across conversations or the average personality of all users.
Approach: They propose a model that conditionally models past conversations to probabilistically model multi-turn conversations in the actor’s persona.
Outcome: The proposed model improves perplexity on 1.7M held out Reddit conversations by 0.47 on scaling from 117M to 8.3B parameters.
Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations (2025.coling-main)

Copied to clipboard

Challenge: Existing LLMs rely on static, predefined personas to capture dynamic and evolving nature of human personalities.
Approach: They propose a dataset with 400,000 conversations and a framework for generating personalized conversations using long-form journal entries from Reddit.
Outcome: The proposed framework generates high-quality, personality-rich dialogues grounded in reddit journal entries.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations