Challenge: Existing dialog-based embodied datasets are not sufficient to develop intelligent navigation-helper agents capable of navigating users in unfamiliar areas.
Approach: They introduce a novel benchmark, Respond to Help Requests, to promote the development of multi-modal navigation helpers capable of responding to requests for help . they also propose two approaches to construct the navigation-helper agent, including fine-tuning a task-oriented multi-mod response generation model that can see and respond, named SeeRee, and employing . a multi-module large language model in a zero-shot manner.
Outcome: The proposed model outperforms the baseline model and the proposed model on two tasks based on human evaluations and automatic benchmarking.

Similar Papers

SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations (2021.emnlp-main)

Copied to clipboard

Challenge: Existing task-oriented dialog datasets do not situate the dialog in the user’s multimodal context.
Approach: They propose to use a dataset to study multimodal task-oriented dialogs in the shopping domain to situate them in the user’s multimodal context.
Outcome: The proposed dataset includes 11K task-oriented user->assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes.
Data Augmentation with Paraphrase Generation and Entity Extraction for Multimodal Dialogue System (2022.lrec-1)

Copied to clipboard

Challenge: Contextually aware intelligent agents are often required to understand the users and their surroundings in real-time.
Approach: They propose to build a multimodal dialogue system for children learning basic math concepts using limited datasets.
Outcome: The proposed system improves the Natural Language Understanding (NLU) module of a task-oriented SDS pipeline with limited dataset resources.
SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR Streams (2023.acl-long)

Copied to clipboard

Challenge: Existing models lack a large-scale benchmark to capture user–assistant interactions . et al., 2022: 145-160.
Approach: They propose a video-grounded task-oriented dialog dataset that captures real-world AI-assisted user scenarios in VR.
Outcome: The proposed dataset captures real-world AI-assisted user scenarios in VR.
Situated and Interactive Multimodal Conversations (2020.coling-main)

Copied to clipboard

Challenge: Situated Interactive MultiModal Conversations (SIMMC) is a new direction for virtual assistants that handle multimodal inputs and perform multimodal actions.
Approach: They propose to use Situated Interactive MultiModal Conversations (SIMMC) to train agents to take multimodal actions grounded in a co-evolving multimodal context.
Outcome: The proposed model will be made publicly available.
Learning a Simple and Effective Model for Multi-turn Response Generation with Auxiliary Tasks (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to multi-turn response generation for open-domain dialogues have a complexity problem . auxiliary tasks that relate to context understanding can guide the learning of the generation model .
Approach: They propose a multi-turn response generation model that has a simple structure yet can effectively leverage conversation contexts for response generation.
Outcome: The proposed model outperforms state-of-the-art models in response quality and human judgment . it also enjoys a faster decoding process .
Learning to Embed Multi-Modal Contexts for Situated Conversational Agents (2022.findings-naacl)

Copied to clipboard

Challenge: Situated Interactive Multi-Modal Conversations 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs.
Approach: They propose a joint learning approach that integrates visual inputs and performs all four subtasks at once for efficiency.
Outcome: The proposed approach won the 10th Dialog Systems Technology Challenge (DSTC10) . it incorporates visual inputs and performs all four subtasks at once for efficiency .
Gated Mechanism Enhanced Multi-Task Learning for Dialog Routing (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for dialog routing are mostly heuristic and cannot achieve high-quality performance.
Approach: They propose a multi-task learning framework with a dialog encoder and two tailored gated mechanism modules to solve this problem.
Outcome: The proposed model can play the role of hierarchical information filtering and is non-invasive to existing dialog systems.
Grounding Task Assistance with Multimodal Cues from a Single Demonstration (2025.findings-acl)

Copied to clipboard

Challenge: RGB video often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences embedded in human behavior.
Approach: They propose a framework that integrates eye gaze and speech cues to improve conversational agents for task assistance by integrating eye gaze with speech cuests.
Outcome: The proposed framework captures fine-grained intent and user-specific cues, enabling richer contextual grounding for visual question answering.
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have shown that virtual agents can help humans achieve task and social goals.
Approach: They propose a tuning-free and label-free method to identify high-quality ICL exemplars for the remediator agent and propose measurable criteria to measure the quality of the negotiation outcomes.
Outcome: The proposed model is able to improve negotiation outcomes across three negotiation topics.
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios is a major challenge for end-to end spoken dialogue models.
Approach: They propose to provide an extensive evaluation framework for end-to-end spoken dialogue models (SDMs) that includes both cognitive dimensions and paralinguistic cues .
Outcome: The proposed benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the spoken dialogue model’s abilities in U**nderstanding, **R**easoning, and **O**ral conversation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations