Challenge: Existing models of reinforcement learning use background planning and may suffer from low-quality simulated experiences.
Approach: They propose a Monte Carlo Tree Search with Double-q Dueling network framework for task-completion dialogue policy learning.
Outcome: The proposed method outperforms the previous model-based reinforcement learning methods and is robust to simulation errors.

Similar Papers

Discriminative Deep Dyna-Q: Robust Planning for Dialogue Policy Learning (D18-1)

Copied to clipboard

Challenge: Existing approaches to improve the effectiveness and robustness of Deep Dyna-Q (DDQ) are based on a discriminator to control the quality of simulated experiences and to improve learning.
Approach: They propose to use an RNN-based discriminator to control the quality of simulated experience to improve the effectiveness and robustness of Deep Dyna-Q.
Outcome: The proposed framework outperforms DDQ by controlling the quality of simulated experience used for planning.
Deep Reinforcement Learning-based Dialogue Policy with Graph Convolutional Q-network (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for deep reinforcement learning lack the ability to learn the relationship between dialogue states and actions.
Approach: They propose a graph-structured dialogue policy framework for task-oriented dialogue systems that uses bipartite graphs to construct two different bipartites and generate user-related and knowledge-related subgraphs.
Outcome: The proposed framework significantly improves the effectiveness and stability of dialogue policies.
Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning (2023.emnlp-main)

Copied to clipboard

Challenge: Optimal policy planning is a difficult task, authors say . many goal-oriented conversations require subjective strategies, they say - a problem in goal-orientated settings .
Approach: They propose an approach to perform goal-oriented dialogue policy planning without model training.
Outcome: The proposed approach performs goal-oriented dialogue policy planning without model training.
Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory Policy (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for training dialogue policies rely on a single learning system, but it requires many rounds of interaction.
Approach: They propose a complementary policy learning framework which exploits the complementary advantages of the episodic memory (EM) policy and the deep Q-network (DQN) policy.
Outcome: The proposed framework outperforms existing methods relying on a single learning system on three dialogue datasets.
Planning Like Human: A Dual-process Framework for Dialogue Planning (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) operate in a reactive mode, often resulting in efficiency issues or suboptimal performance.
Approach: They propose a dual-process dialogue planning framework that leverages the dual-process theory of human cognition and a deliberative Monte Carlo Tree Search mechanism to emulate human-like conversational dynamics.
Outcome: The proposed framework outperforms existing methods in achieving high-quality dialogues and operational efficiency.
Few-Shot Structured Policy Learning for Multi-Domain and Multi-Task Dialogues (2023.findings-eacl)

Copied to clipboard

Challenge: Reinforcement learning is widely adopted to model dialogue managers in task-oriented dialogues, but the user simulator provided by state-of-the-art dialogue frameworks are only rough approximations of human behaviour.
Approach: They propose to use structured policies to improve sample efficiency when learning on multi-domain and multi-task environments.
Outcome: The proposed policies improve sample efficiency and performance on multi-domain and multi-task environments.
Deep Dyna-Q: Integrating Planning for Task-Completion Dialogue Policy Learning (P18-1)

Copied to clipboard

Challenge: Training a task-completion dialogue agent via reinforcement learning (RL) is costly because it requires many interactions with real users.
Approach: They propose a framework that integrates planning for task-completion dialogue policy learning into a dialogue agent using a world model to mimic real user response and generate simulated experience.
Outcome: The proposed framework integrates planning for task-completion dialogue policy learning with real user interaction and simulated user behavior.
Budgeted Policy Learning for Task-Oriented Dialogue Systems (P19-1)

Copied to clipboard

Challenge: Existing approaches to integrating reinforcement learning into task-oriented dialogue systems require a fixed, small amount of user interactions to learn.
Approach: They propose a budget-conscious scheduling approach that optimizes a fixed, small amount of user interactions for dialogue agent learning.
Outcome: The proposed approach improves on a movie-ticket booking task with simulated and real users.
Subgoal Discovery for Hierarchical Dialogue Policy Learning (D18-1)

Copied to clipboard

Challenge: Existing methods to develop dialogue agents for complex tasks require sparse reward signals.
Approach: They propose a divide-and-conquer approach that exploits the hidden structure of a task . they use subgoals to divide a goal-oriented task into simpler subgoal sets .
Outcome: The proposed approach performs competitively against state-of-the-art methods that require human-defined subgoals.
Rethinking Supervised Learning and Reinforcement Learning in Task-Oriented Dialogue Systems (2020.findings-emnlp)

Copied to clipboard

Challenge: Dialogue policy learning for task-oriented dialogue systems has enjoyed great progress through using reinforcement learning methods.
Approach: They propose a dialogue action decoder and a simulator-free adversarial learning method to improve dialogue agent performance without using reinforcement learning.
Outcome: The proposed methods achieve more stable and higher performance with fewer efforts, such as the domain knowledge required to design a user simulator and the intractable parameter tuning in reinforcement learning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations