Multimodal Dialogue State Tracking (2022.naacl-main)

Copied to clipboard

Challenge: Dialogue state tracking is a key component of dialogue systems.
Approach: They propose to extend the definition of dialogue state tracking to multimodality . they propose a new synthetic benchmark and a novel baseline for this task .
Outcome: The proposed task is based on a synthetic benchmark and a self-supervised video understanding task.

Similar Papers

Scaling Multi-Domain Dialogue State Tracking via Query Reformulation (N19-2)

Copied to clipboard

Challenge: Using a pointer-generator network, we model the reference resolution task as a dialogue context-aware user query reformulation task.
Approach: They propose a pointer-generator network and a novel multi-task learning setup to model dialogue state tracking and referring expression resolution tasks using a dialogue context-aware user query reformulation task.
Outcome: The proposed model improves absolute F1 on internal and public benchmarks.
Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue Systems (P19-1)

Copied to clipboard

Challenge: Existing work on video-grounded dialogue systems is limited by feature space and semantic information.
Approach: They propose multimodal transformer networks to encode videos and incorporate information from different modalities.
Outcome: The proposed system generates appropriate conversational response to queries of humans based on visual and audio aspects of a given video . it also generalizes to another multimodal visual-grounded dialogue task, and obtains promising performance.
Dialogue State Tracking with Incremental Reasoning (2021.tacl-1)

Copied to clipboard

Challenge: Empirical results show that our method outperforms the state-of-the-art methods in terms of joint belief accuracy.
Approach: They propose to track dialogue states gradually with reasoning over dialogue turns using the back-end data.
Outcome: Empirical results show that the proposed method outperforms state-of-the-art methods in terms of joint belief accuracy for a large-scale human–human dialogue dataset.
Scalable and Accurate Dialogue State Tracking via Hierarchical Sequence Generation (D19-1)

Copied to clipboard

Challenge: Existing approaches to dialogue state tracking rely on pre-defined ontologies . however, these methods suffer from computational complexity that increases proportionally to the number of pre-determined slots.
Approach: They propose a model that generates a sequence of belief states without the pre-defined ontology list.
Outcome: The proposed model scales easily with the increasing number of pre-defined slots and domains and reaches the state-of-the-art performance on the multi-domain and single domain dialogue state tracking datasets.
Towards Universal Dialogue State Tracking (D18-1)

Copied to clipboard

Challenge: Existing approaches to dialogue state tracking are difficult to scale to large dialogue domains.
Approach: They propose a universal dialogue state tracker that is independent of the number of values and shares parameters across all slots.
Outcome: The proposed system significantly outperforms state-of-the-art approaches on two datasets.
Beyond the Granularity: Multi-Perspective Dialogue Collaborative Selection for Dialogue State Tracking (2022.acl-long)

Copied to clipboard

Challenge: Experimental results show that task-oriented dialogue systems have attracted growing attention and achieved substantial progress.
Approach: They propose a method that dynamically selects relevant dialogue contents for each slot . they retrieve turn-level utterances and evaluate their relevance to the slot from three perspectives .
Outcome: The proposed method achieves state-of-the-art performance on MultiWOZ 2.1 and MultiWOz 2.2 and superior performance on multiple mainstream benchmark datasets.
Multi-Domain Dialogue State Tracking By Neural-Retrieval Augmentation (2022.findings-aacl)

Copied to clipboard

Challenge: Existing approaches for DST are conditioned on previous dialogue states, but the dependency on previous dialogs makes it difficult to prevent error propagation to subsequent turns.
Approach: They propose to create a Neural Index based on dialogue context by analyzing user dialogue and previous turn state and generating a retrieval-guided generation approach.
Outcome: The proposed framework retrieves dialogue context from the index built using unstructured dialogue state and structured user/system utterances.
A Sequence-to-Sequence Approach to Dialogue State Tracking (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for dialogue state tracking are still challenging, but they are improving . a new approach to dialogue state monitoring is proposed, called Seq2Seq-DU .
Approach: They propose a new dialogue state tracking module that formalizes DST as a sequence-to-sequence problem.
Outcome: The proposed method outperforms existing methods on benchmark datasets in different settings.
Ordinal and Attribute Aware Response Generation in a Multimodal Dialogue System (P19-1)

Copied to clipboard

Challenge: Existing multimodal dialogue systems are based on unimodal sources, capturing information from text and image.
Approach: They propose a position and attribute aware attention mechanism to learn enhanced image representation conditioned on the user utterance.
Outcome: The proposed model outperforms the state-of-the-art models on text similarity metrics.
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking (2021.emnlp-main)

Copied to clipboard

Challenge: Existing models for dialogue state tracking are based on Graph Attention Networks . if the relationship between slots and values is modelled explicitly, this can be improved .
Approach: They propose a model architecture that augments GPT-2 with Graph Attention Networks to allow sequential prediction of slot values.
Outcome: The proposed architecture improves performance against a strong GPT-2 baseline and with sparsely supervised training.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations