A Visually-grounded First-person Dialogue Dataset with Verbal and Non-verbal Responses (2020.emnlp-main)
Copied to clipboard
| Challenge: | In visual-grounded dialogue systems, first-person visual information about where the other speakers are and what they are paying attention to is crucial to understand their intentions. |
| Approach: | They propose a visually-grounded first-person dialogue (VFD) dataset with verbal and non-verbal responses. |
| Outcome: | The proposed dataset provides verbal and non-verbal responses for first-person visual information and recent neural network models. |
Similar Papers
A Linguistic Analysis of Visually Grounded Dialogues Based on Spatial Expressions (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models for visually grounded dialogues often contain undesirable biases and lack sophisticated linguistic analyses, making it difficult to understand how well they recognize their precise linguistic structures. |
| Approach: | They propose a framework for studying fine-grained language understanding in visually grounded dialogues by using a common grounding dataset which contains minimal bias by design. |
| Outcome: | The proposed framework can reveal both strengths and weaknesses of baseline models in essential levels of detail. |
Game-Based Video-Context Dialogue (D18-1)
Copied to clipboard
| Challenge: | Current dialogue systems focus more on textual and speech context knowledge and are usually based on two speakers. |
| Approach: | They propose to use live soccer game videos and Twitch.tv chats to develop visual-grounded dialogue models. |
| Outcome: | The proposed model can generate relevant temporal and spatial event language from live video and chat history while also being relevant to chat history. |
Image-Chat: Engaging Grounded Conversations (2020.acl-main)
Copied to clipboard
| Challenge: | In order for machines to communicate with humans, they must understand the natural things that humans say about the world they live in and respond in kind. |
| Approach: | They propose to fuse a set of neural architectures using image and text representations to achieve this goal. |
| Outcome: | The proposed model performs well on the Image-Chat task and humans prefer it 47.7% of the time. |
VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic Transitions (2023.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for video-grounded dialogues neglect the intrinsic attributes of multimodal dialogues, such as scene and topic transitions. |
| Approach: | They propose to use a large scale video-grounded scene&topic AwaRe dialogue dataset to study video-based dialogue understanding. |
| Outcome: | The proposed dataset shows that multimodal information and segments are important in video-grounded dialogue understanding and generation. |
Textual Supervision for Visually Grounded Spoken Language Understanding (2020.findings-emnlp)
Copied to clipboard
| Challenge: | a new approach to spoken language understanding extracts semantic information directly from speech without relying on transcriptions. |
| Approach: | They propose to use textual supervision to train visually-grounded models of spoken language understanding without relying on transcriptions. |
| Outcome: | The proposed model improves when enough text is available, the study shows . compared with pipeline-based models, the pipeline approach performs better when enough data is available . |
The PhotoBook Dataset: Building Common Ground through Visually-Grounded Dialogue (P19-1)
Copied to clipboard
| Challenge: | Using the PhotoBook dataset, we investigate shared dialogue history accumulating during conversation . human interlocutors are known to collaboratively establish a shared repository of mutual information during a conversation - this common ground is then used to optimise understanding and communication efficiency. |
| Approach: | They propose a data-collection task formulated as a collaborative game prompting two online participants to refer to images utilising both their visual context and previously established referring expressions. |
| Outcome: | The proposed model takes into account shared information accumulated in a reference chain and is important to resolve later descriptions. |
doc2dial: A Goal-Oriented Document-Grounded Dialogue Dataset (2020.emnlp-main)
Copied to clipboard
| Challenge: | doc2dial dataset is a goal-oriented document-grounded dialogue model . it is based on how the authors compose documents for guiding end users . |
| Approach: | They propose a dataset of goal-oriented dialogues grounded in documents . they use annotated conversations with an average of 14 turns to generate conversational utterances . |
| Outcome: | The proposed dataset includes over 4500 annotated conversations with an average of 14 turns grounded in over 450 documents from four domains. |
DIRECT: Direct and Indirect Responses in Conversational Text Corpus (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Neural conversation models have been able to generate fluent responses through training on a dialogue corpus, but they lack the ability to reveal the implied intentions of users. |
| Approach: | They propose to train neural conversation models on a dialogue corpus that provides pragmatic paraphrases to advance techniques for natural language understanding in dialogue systems. |
| Outcome: | The proposed corpus provides 71,498 pairs of indirect–direct utterance pairs accompanied by a multi-turn dialogue history extracted from the MultiWoZ dataset. |
What You See is What You Get: Visual Pronoun Coreference Resolution in Dialogues (D19-1)
Copied to clipboard
| Challenge: | a core task of natural language understanding is to ground a pronoun to a visual object it refers to . problem arises when people use pronounos to refer to something they can see without prior introduction . a novel visual-aware PCR model is proposed to solve this problem . |
| Approach: | They propose a visual-aware PCR model to ground a pronoun to a visible object . they propose PCR using a large-scale dialogue dataset to investigate this problem . |
| Outcome: | The proposed model can help resolve pronouns in conversational contexts. |
DialogSum: A Real-Life Scenario Dialogue Summarization Dataset (2021.findings-acl)
Copied to clipboard
| Challenge: | Experimental results show unique challenges in dialogue summarization such as spoken terms, special discourse structures, coreferences and ellipsis, pragmatics and social common sense. |
| Approach: | They propose a large-scale labeled dialogue summarization dataset . they use state-of-the-art neural models to analyze spoken dialogue summaries . |
| Outcome: | The proposed dataset can be used to analyze spoken dialogue summarization challenges. |