| Challenge: | Recent advances focus on improving DRL-based dialogue policy optimization. |
| Approach: | They propose to design a graph neural network structure that is better suited for dialogue management. |
| Outcome: | The proposed approach outperforms state-of-the-art approaches in 18 tasks of the PyDial benchmark. |
Similar Papers
Deep Reinforcement Learning-based Dialogue Policy with Graph Convolutional Q-network (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for deep reinforcement learning lack the ability to learn the relationship between dialogue states and actions. |
| Approach: | They propose a graph-structured dialogue policy framework for task-oriented dialogue systems that uses bipartite graphs to construct two different bipartites and generate user-related and knowledge-related subgraphs. |
| Outcome: | The proposed framework significantly improves the effectiveness and stability of dialogue policies. |
Few-Shot Structured Policy Learning for Multi-Domain and Multi-Task Dialogues (2023.findings-eacl)
Copied to clipboard
| Challenge: | Reinforcement learning is widely adopted to model dialogue managers in task-oriented dialogues, but the user simulator provided by state-of-the-art dialogue frameworks are only rough approximations of human behaviour. |
| Approach: | They propose to use structured policies to improve sample efficiency when learning on multi-domain and multi-task environments. |
| Outcome: | The proposed policies improve sample efficiency and performance on multi-domain and multi-task environments. |
Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory Policy (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for training dialogue policies rely on a single learning system, but it requires many rounds of interaction. |
| Approach: | They propose a complementary policy learning framework which exploits the complementary advantages of the episodic memory (EM) policy and the deep Q-network (DQN) policy. |
| Outcome: | The proposed framework outperforms existing methods relying on a single learning system on three dialogue datasets. |
Rethinking Supervised Learning and Reinforcement Learning in Task-Oriented Dialogue Systems (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Dialogue policy learning for task-oriented dialogue systems has enjoyed great progress through using reinforcement learning methods. |
| Approach: | They propose a dialogue action decoder and a simulator-free adversarial learning method to improve dialogue agent performance without using reinforcement learning. |
| Outcome: | The proposed methods achieve more stable and higher performance with fewer efforts, such as the domain knowledge required to design a user simulator and the intractable parameter tuning in reinforcement learning. |
Feudal Reinforcement Learning for Dialogue Management in Large Domains (N18-2)
Copied to clipboard
Iñigo Casanueva, Paweł Budzianowski, Pei-Hao Su, Stefan Ultes, Lina M. Rojas-Barahona, Bo-Hsiang Tseng, Milica Gašić
| Challenge: | Reinforcement learning (RL) is a promising approach to model dialogue policy optimisation but fails to scale to large domains due to the curse of dimensionality. |
| Approach: | They propose a novel approach to dialogue policy optimisation using reinforcement learning . they propose to decompose the decision into two steps using a domain ontology . |
| Outcome: | The proposed architecture outperforms state-of-the-art in several dialogue domains without any additional reward signal. |
Deep Dyna-Q: Integrating Planning for Task-Completion Dialogue Policy Learning (P18-1)
Copied to clipboard
| Challenge: | Training a task-completion dialogue agent via reinforcement learning (RL) is costly because it requires many interactions with real users. |
| Approach: | They propose a framework that integrates planning for task-completion dialogue policy learning into a dialogue agent using a world model to mimic real user response and generate simulated experience. |
| Outcome: | The proposed framework integrates planning for task-completion dialogue policy learning with real user interaction and simulated user behavior. |
An Efficient Dialogue Policy Agent with Model-Based Causal Reinforcement Learning (2025.coling-main)
Copied to clipboard
| Challenge: | Existing models for dialogue policy training consider one-step dialogues, leading to inaccurate simulations. |
| Approach: | They propose a framework for dialogue policy learning that trains an agent to select dialogue actions via deep reinforcement learning. |
| Outcome: | The proposed framework achieves state-of-the-art performance on three dialogue datasets . it uses model-based reinforcement learning with automatically constructed causal chains . |
Gaussian Process based Deep Dyna-Q approach for Dialogue Policy Learning (2021.findings-acl)
Copied to clipboard
| Challenge: | Reinforcement learning (RL) is the main dialogue policy learning method in recent years. |
| Approach: | They propose a Gaussian Process based Deep Dyna-Q approach to dialogue policy learning . they propose evaluating the quality of experiences generated by the world model using a discriminator . |
| Outcome: | The proposed approach improves the effectiveness and efficiency of dialogue policy learning by 20% with fewer human-machine interactions. |
Deep Reinforcement Learning for NLP (P18-5)
Copied to clipboard
| Challenge: | Many natural language processing tasks can be formulated as deep reinforcement learning (DRL) problems. |
| Approach: | This tutorial provides an introduction to the foundations of deep reinforcement learning . it describes recent advances in designing deep reinforcement for NLP . |
| Outcome: | This tutorial provides an introduction to the foundations of deep reinforcement learning and some practical solutions for NLP tasks. |
Deep Reinforcement Learning with Hierarchical Action Exploration for Dialogue Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to improve dialogues with random sampling are inefficient due to the large number of eligible responses with high action values. |
| Approach: | They propose a dual-granularity Q-function that extracts actions based on a grained hierarchy . they use offline RL and learn from multiple reward functions designed to capture emotional nuances in human interactions. |
| Outcome: | The proposed approach outperforms baselines across automatic metrics and human evaluations. |