Challenge: Existing embodied navigation methods struggle with such tasks due to their limitations in comprehending high-level human instructions and localizing objects with an open vocabulary.
Approach: They propose a hierarchical framework for long-horizon navigation that integrates human instructions with 3D scene views.
Outcome: The proposed model achieves SOTA results and can complete long-horizon navigation tasks across different robot embodiments in real-world environments.

Similar Papers

NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM (2025.findings-acl)

Copied to clipboard

Challenge: High-performance vision-and-language navigation models require large amounts of training data, the high cost of manual annotating has seriously hindered this field.
Approach: They propose a retrieval-augmented generation framework that generates user demand instructions for vision-and-language navigation.
Outcome: The proposed model achieves SOTA performance on the REVERIE benchmark.
ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments (2020.findings-emnlp)

Copied to clipboard

Challenge: embodied agents are expected to perform specific tasks after reaching the destination . a novel vision-and-language navigation task is designed to support this task .
Approach: They combine vision-and-language navigation, assembling objects and object referring expression comprehension to create a joint navigation-and assembly task.
Outcome: The proposed task is based on vision-and-language navigation and assembly . it uses human-written navigation and assembling instructions and ground truth trajectories . the large model-human performance gap shows that the task is challenging and wide scope for future work.
NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and training paradigms.
Approach: They propose a visual-target trajectory collection pipeline that generates trajectories for GUI and embodied tasks using a single formulation.
Outcome: The proposed agent outperforms state-of-the-art agents in GUI navigation, spatial affordance prediction, and embodied navigation.
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing speaker models learn strategies to evade evaluation metrics and obtain higher scores even for low-quality sentences.
Approach: They propose a speaker-based instruction generator that utilises both structural and semantic knowledge of the environment to produce richer instructions.
Outcome: The proposed model outperforms existing models and is evaluated using standard metrics.
GesNavi: Gesture-guided Outdoor Vision-and-Language Navigation (2024.eacl-srw)

Copied to clipboard

Challenge: Existing datasets for outdoor Vision-and-Language Navigation (VLN) tasks do not include verbal instructions for communicating with mobility.
Approach: They propose a dataset for gesture-guided outdoor VLN instructions with demonstrative expressions that incorporates gestures and linguistic commands.
Outcome: The proposed datasets are compared against existing datasets and analysed in detail.
Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments (2021.emnlp-main)

Copied to clipboard

Challenge: Prior work on Vision-and-Language Navigation (VLN) tasks do not measure how much of a language instruction the agent is able to follow.
Approach: They propose a language-aligned supervision scheme that measures the number of sub-instructions the agent has completed during navigation.
Outcome: The proposed method is based on the previous work on the Vision-and-Language Navigation task, which assumes a discrete navigation graph (navgraph) but not on the current work.
Into the Unknown: Generating Geospatial Descriptions for New Environments (2024.findings-acl)

Copied to clipboard

Challenge: Similar to vision-and-language navigation tasks, the Rendezvous (RVS) task requires reasoning over allocentric spatial relationships using non-sequential navigation instructions and maps.
Approach: They propose a large-scale augmentation method for generating high-quality synthetic data for new environments using readily available geospatial data.
Outcome: The proposed method improves accuracy on unseen and seen environments by 45.83% on the Rendezvous (RVS) task.
Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions (2022.acl-long)

Copied to clipboard

Challenge: Vision-and-Language Navigation (VLN) is a research topic that is gaining attention in the field of artificial intelligence.
Approach: They propose to build an embodied agent that can communicate with humans in natural language and navigate in real 3D environments.
Outcome: This paper reviews current studies in the emerging field of vision-and-language navigation . it highlights limitations and opportunities for future work .
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models (2024.acl-short)

Copied to clipboard

Challenge: Recent studies have revealed significant deficiencies of LVLMs in understanding visual contents, leaving the gap between current embodied intelligence and large vision-language models (LVLM) .
Approach: They propose to use a benchmark to evaluate LVLMs' spatial understanding of embodied environments to evaluate their ability to understand visual contents.
Outcome: The proposed benchmark is derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective.
Can MLLMs Find Their Way in a City? Exploring Emergent Navigation from Web-Scale Knowledge (2026.eacl-long)

Copied to clipboard

Challenge: Existing evaluation benchmarks for multimodal large language models (MLLMs) are language-centric or heavily reliant on simulated environments, rarely probing the nuanced, knowledge-intensive reasoning essential for practical, real-world scenarios.
Approach: They propose a task of Sparsely Grounded Visual Navigation to evaluate MLLM-driven agents in city navigation in four diverse global cities.
Outcome: The proposed benchmark encompassing four diverse global cities evaluates agents' decision-making abilities in city navigation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations