The PhotoBook Dataset: Building Common Ground through Visually-Grounded Dialogue (P19-1)
Copied to clipboard
| Challenge: | Using the PhotoBook dataset, we investigate shared dialogue history accumulating during conversation . human interlocutors are known to collaboratively establish a shared repository of mutual information during a conversation - this common ground is then used to optimise understanding and communication efficiency. |
| Approach: | They propose a data-collection task formulated as a collaborative game prompting two online participants to refer to images utilising both their visual context and previously established referring expressions. |
| Outcome: | The proposed model takes into account shared information accumulated in a reference chain and is important to resolve later descriptions. |
Similar Papers
Illustrative Language Understanding: Large-Scale Visual Grounding with Image Search (P18-1)
Copied to clipboard
| Challenge: | a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search is currently used to learn word representations. |
| Approach: | They propose a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search. |
| Outcome: | The proposed model is based on a large-scale lookup operation to ground language using image search. |
doc2dial: A Goal-Oriented Document-Grounded Dialogue Dataset (2020.emnlp-main)
Copied to clipboard
| Challenge: | doc2dial dataset is a goal-oriented document-grounded dialogue model . it is based on how the authors compose documents for guiding end users . |
| Approach: | They propose a dataset of goal-oriented dialogues grounded in documents . they use annotated conversations with an average of 14 turns to generate conversational utterances . |
| Outcome: | The proposed dataset includes over 4500 annotated conversations with an average of 14 turns grounded in over 450 documents from four domains. |
Collecting Visually-Grounded Dialogue with A Game Of Sorts (2022.lrec-1)
Copied to clipboard
| Challenge: | referring in conversation is a collaborative process that cannot be described as an exchange of minimally-specified referring expressions. |
| Approach: | They propose a collaborative image ranking task that allows players to agree on a sorting criteria. |
| Outcome: | The proposed game aims to ground referring expressions in visually-grounded dialogues . it uses a game-like approach to rank images in a role-symmetric dialogue . |
Game-Based Video-Context Dialogue (D18-1)
Copied to clipboard
| Challenge: | Current dialogue systems focus more on textual and speech context knowledge and are usually based on two speakers. |
| Approach: | They propose to use live soccer game videos and Twitch.tv chats to develop visual-grounded dialogue models. |
| Outcome: | The proposed model can generate relevant temporal and spatial event language from live video and chat history while also being relevant to chat history. |
Image-Chat: Engaging Grounded Conversations (2020.acl-main)
Copied to clipboard
| Challenge: | In order for machines to communicate with humans, they must understand the natural things that humans say about the world they live in and respond in kind. |
| Approach: | They propose to fuse a set of neural architectures using image and text representations to achieve this goal. |
| Outcome: | The proposed model performs well on the Image-Chat task and humans prefer it 47.7% of the time. |
A Linguistic Analysis of Visually Grounded Dialogues Based on Spatial Expressions (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models for visually grounded dialogues often contain undesirable biases and lack sophisticated linguistic analyses, making it difficult to understand how well they recognize their precise linguistic structures. |
| Approach: | They propose a framework for studying fine-grained language understanding in visually grounded dialogues by using a common grounding dataset which contains minimal bias by design. |
| Outcome: | The proposed framework can reveal both strengths and weaknesses of baseline models in essential levels of detail. |
A Corpus for Reasoning about Natural Language Grounded in Photographs (P19-1)
Copied to clipboard
| Challenge: | a dataset for visual reasoning with natural language and images is available. |
| Approach: | They propose a dataset for joint reasoning about natural language and images . they crowdsource 107,292 examples of English sentences paired with web photographs . |
| Outcome: | The proposed dataset combines 107,292 examples of English sentences with web photographs . Qualitative analysis shows the data requires compositional joint reasoning . |
Dialogue Collection for Recording the Process of Building Common Ground in a Collaborative Task (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on the process of building common ground have not been well conducted. |
| Approach: | They propose a method for recording the process of building common ground through a dialogue by using the intermediate result of a task. |
| Outcome: | The proposed method can record the building common ground process by using the intermediate result of a task and can be estimated quite accurately. |
DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogue (2021.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks do not have enough annotations to analyze video-grounded dialogue systems and understand their capabilities and limitations in isolation. |
| Approach: | They present a Diagnostic Dataset for Video-grounded dialogue with minimal biases and detailed annotations for the different types of reasoning over the spatio-temporal space of video. |
| Outcome: | The proposed system is based on 11k CATER synthetic videos and contains 10 instances of 10-round dialogues for each video. |
A Dataset for Document Grounded Conversations (D18-1)
Copied to clipboard
| Challenge: | a dataset of document grounded conversations provides information on content of a document . current datasets lacking conversation grounding do not provide this information . |
| Approach: | They propose a document grounded dataset for conversations . they use Wikipedia articles about popular movies to define document grounded conversations based on their results . |
| Outcome: | The proposed dataset provides a source of information and provides benchmark performance on the task of generating the next response. |