Challenge: Using the PhotoBook dataset, we investigate shared dialogue history accumulating during conversation . human interlocutors are known to collaboratively establish a shared repository of mutual information during a conversation - this common ground is then used to optimise understanding and communication efficiency.
Approach: They propose a data-collection task formulated as a collaborative game prompting two online participants to refer to images utilising both their visual context and previously established referring expressions.
Outcome: The proposed model takes into account shared information accumulated in a reference chain and is important to resolve later descriptions.

Similar Papers

Illustrative Language Understanding: Large-Scale Visual Grounding with Image Search (P18-1)

Copied to clipboard

Challenge: a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search is currently used to learn word representations.
Approach: They propose a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search.
Outcome: The proposed model is based on a large-scale lookup operation to ground language using image search.
doc2dial: A Goal-Oriented Document-Grounded Dialogue Dataset (2020.emnlp-main)

Copied to clipboard

Challenge: doc2dial dataset is a goal-oriented document-grounded dialogue model . it is based on how the authors compose documents for guiding end users .
Approach: They propose a dataset of goal-oriented dialogues grounded in documents . they use annotated conversations with an average of 14 turns to generate conversational utterances .
Outcome: The proposed dataset includes over 4500 annotated conversations with an average of 14 turns grounded in over 450 documents from four domains.
Collecting Visually-Grounded Dialogue with A Game Of Sorts (2022.lrec-1)

Copied to clipboard

Challenge: referring in conversation is a collaborative process that cannot be described as an exchange of minimally-specified referring expressions.
Approach: They propose a collaborative image ranking task that allows players to agree on a sorting criteria.
Outcome: The proposed game aims to ground referring expressions in visually-grounded dialogues . it uses a game-like approach to rank images in a role-symmetric dialogue .
Game-Based Video-Context Dialogue (D18-1)

Copied to clipboard

Challenge: Current dialogue systems focus more on textual and speech context knowledge and are usually based on two speakers.
Approach: They propose to use live soccer game videos and Twitch.tv chats to develop visual-grounded dialogue models.
Outcome: The proposed model can generate relevant temporal and spatial event language from live video and chat history while also being relevant to chat history.
Image-Chat: Engaging Grounded Conversations (2020.acl-main)

Copied to clipboard

Challenge: In order for machines to communicate with humans, they must understand the natural things that humans say about the world they live in and respond in kind.
Approach: They propose to fuse a set of neural architectures using image and text representations to achieve this goal.
Outcome: The proposed model performs well on the Image-Chat task and humans prefer it 47.7% of the time.
A Linguistic Analysis of Visually Grounded Dialogues Based on Spatial Expressions (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models for visually grounded dialogues often contain undesirable biases and lack sophisticated linguistic analyses, making it difficult to understand how well they recognize their precise linguistic structures.
Approach: They propose a framework for studying fine-grained language understanding in visually grounded dialogues by using a common grounding dataset which contains minimal bias by design.
Outcome: The proposed framework can reveal both strengths and weaknesses of baseline models in essential levels of detail.
A Corpus for Reasoning about Natural Language Grounded in Photographs (P19-1)

Copied to clipboard

Challenge: a dataset for visual reasoning with natural language and images is available.
Approach: They propose a dataset for joint reasoning about natural language and images . they crowdsource 107,292 examples of English sentences paired with web photographs .
Outcome: The proposed dataset combines 107,292 examples of English sentences with web photographs . Qualitative analysis shows the data requires compositional joint reasoning .
Dialogue Collection for Recording the Process of Building Common Ground in a Collaborative Task (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on the process of building common ground have not been well conducted.
Approach: They propose a method for recording the process of building common ground through a dialogue by using the intermediate result of a task.
Outcome: The proposed method can record the building common ground process by using the intermediate result of a task and can be estimated quite accurately.
DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogue (2021.acl-long)

Copied to clipboard

Challenge: Existing benchmarks do not have enough annotations to analyze video-grounded dialogue systems and understand their capabilities and limitations in isolation.
Approach: They present a Diagnostic Dataset for Video-grounded dialogue with minimal biases and detailed annotations for the different types of reasoning over the spatio-temporal space of video.
Outcome: The proposed system is based on 11k CATER synthetic videos and contains 10 instances of 10-round dialogues for each video.
A Dataset for Document Grounded Conversations (D18-1)

Copied to clipboard

Challenge: a dataset of document grounded conversations provides information on content of a document . current datasets lacking conversation grounding do not provide this information .
Approach: They propose a document grounded dataset for conversations . they use Wikipedia articles about popular movies to define document grounded conversations based on their results .
Outcome: The proposed dataset provides a source of information and provides benchmark performance on the task of generating the next response.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations