Challenge: a goal-driven collaborative drawing task combines language, perception, and actions in a partially observable environment . et al., 1990: 138K messages exchanged between human players.
Approach: They propose a goal-driven collaborative task that combines language, perception, and action . they collect a clip art dataset and use it to build an image-drawing game between two agents .
Outcome: The proposed task integrates language, perception, and action in a virtual world . it is based on a dataset of 10K dialogs and 138K messages exchanged between humans .

Similar Papers

CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning (2020.acl-main)

Copied to clipboard

Challenge: Approaches to Grounded Language Learning focus on a single task-based final performance measure which may not depend on desirable properties of the learned hidden representations.
Approach: They propose an evaluation framework for Grounded Language Learning with Attributes based on three sub-tasks: 1) Goal-oriented evaluation; 2) Object attribute prediction evaluation; and 3) Zero-shot evaluation.
Outcome: The proposed framework evaluates the quality of learned representations with respect to attribute grounding.
Grounding Language in Multi-Perspective Referential Communication (2024.emnlp-main)

Copied to clipboard

Challenge: Using a dataset of 2,970 human-written referring expressions, we find that the performance of automated models in both reference generation and comprehension lags behind that of pairs of human agents.
Approach: They propose a task and dataset for referring expression generation and comprehension in multi-agent embodied environments where two agents must take into account one another's visual perspective to produce and understand references to objects in a scene.
Outcome: The proposed model outperforms the strongest proprietary model and improves communicative success from 58.9 to 69.3% when trained with a listener.
Image-Chat: Engaging Grounded Conversations (2020.acl-main)

Copied to clipboard

Challenge: In order for machines to communicate with humans, they must understand the natural things that humans say about the world they live in and respond in kind.
Approach: They propose to fuse a set of neural architectures using image and text representations to achieve this goal.
Outcome: The proposed model performs well on the Image-Chat task and humans prefer it 47.7% of the time.
Instruction Clarification Requests in Multimodal Collaborative Dialogue Games: Tasks, and an Analysis of the CoDraw Dataset (2023.eacl-main)

Copied to clipboard

Challenge: In visual instruction-following dialogue games, players can engage in repair mechanisms in the face of an ambiguous or underspecified instruction.
Approach: They annotate Instruction Clarification Requests (iCRs) in CoDraw, an existing dataset of interactions in a multimodal collaborative dialogue game.
Outcome: The proposed dataset contains lexically and semantically diverse iCRs produced self-motivatedly by players deciding to clarify in order to solve the task successfully.
Paparazzi: A Deep Dive into the Capabilities of Language and Vision Models for Grounding Viewpoint Descriptions (2023.findings-eacl)

Copied to clipboard

Challenge: Existing language and vision models can be used for language understanding in 3D environments . however, existing models lack specific properties and biases that limit their performance .
Approach: They propose a framework that uses a camera to generate images from different viewpoints and evaluate them in terms of their similarity to natural language descriptions.
Outcome: The proposed model performs poorly on most canonical views and fine-tunes using hard negative sampling and random contrasting yields good results even under conditions with little available training data.
TuringAdvice: A Generative and Dynamic Evaluation of Language Use (2021.naacl-main)

Copied to clipboard

Challenge: Empirical results show that today’s language models struggle at TuringAdvice . language models are getting ever-larger, and are being trained on ever-increasing quantities of text .
Approach: They propose a task task that requires models to generate helpful advice in natural language.
Outcome: The proposed model outperforms even multibillion parameter models on 600k in-domain training examples.
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback (2024.findings-eacl)

Copied to clipboard

Challenge: In many approaches to Natural Language Processing tasks, language is inherently interactive.
Approach: They propose to use human-AI collaboration to improve human-human interaction by providing feedback that the agent can understand and utilize.
Outcome: The proposed task is an interactive grounded language understanding task in a MineCraft-like world.
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for visual grounding rely on the assumption that the given expression must be literal . this impedes the practical deployment of agents in real-world scenarios.
Approach: They propose a visual grounding task that uses intention expressions to locate foreground entities . they build a large-scale IVG dataset with free-form intention expression to promote VG .
Outcome: The proposed method is based on a large-scale intention-driven visual-language (V-L) dataset with free-form intention expressions.
The Dialogue Dodecathlon: Open-Domain Knowledge and Image Grounded Conversational Agents (2020.acl-main)

Copied to clipboard

Challenge: a set of 12 tasks that measure if a conversational agent can communicate engagingly with personality and empathy, ask questions, answer questions by utilizing knowledge resources, and perceive and converse about images.
Approach: They propose a set of 12 tasks that measure if a conversational agent can communicate engagingly with personality and empathy . they use large dialogue datasets to multi-task and obtain state-of-the-art results .
Outcome: The proposed model improves over a BERT pre-trained model on large dialogue datasets and provides state-of-the-art results on many of the tasks.
Collecting Visually-Grounded Dialogue with A Game Of Sorts (2022.lrec-1)

Copied to clipboard

Challenge: referring in conversation is a collaborative process that cannot be described as an exchange of minimally-specified referring expressions.
Approach: They propose a collaborative image ranking task that allows players to agree on a sorting criteria.
Outcome: The proposed game aims to ground referring expressions in visually-grounded dialogues . it uses a game-like approach to rank images in a role-symmetric dialogue .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations