Challenge: We observe two kinds of instructions that make the grounding in the vision-and-language navigation task quite challenging.
Approach: They propose to use a translator module to convert instructions into easy-to-follow sub-instruction representations at each step.
Outcome: The proposed model is based on a Room2Room (R2R), Room4room (R4R), and Room2room Last (R1R-Last) datasets and achieves state-of-the-art results on multiple benchmarks.

Similar Papers

Sub-Instruction Aware Vision-and-Language Navigation (2020.emnlp-main)

Copied to clipboard

Challenge: Despite significant advances, few previous works are able to fully utilize the strong correspondence between visual and textual sequences.
Approach: They propose to provide agents with fine-grained annotations during training and provide them with sub-instructions and their corresponding paths.
Outcome: The proposed method improves the performance of four state-of-the-art agents in a room-to-room (R2R) benchmark dataset.
CLEAR: Improving Vision-Language Navigation with Cross-Lingual, Environment-Agnostic Representations (2022.findings-naacl)

Copied to clipboard

Challenge: Using multilingual instructions to learn a better cross-lingual representation is challenging for multilingual agents.
Approach: They propose to use multilingual instructions to learn a shared cross-lingual language representation for the three languages in a Room-Across-Room dataset.
Outcome: The proposed model improves on the room-Across-room and vision-and-dialogue navigation tasks by maximizing similarity between semantically aligned image pairs from different environments.
Diagnosing Vision-and-Language Navigation: What Really Matters (2022.naacl-main)

Copied to clipboard

Challenge: Existing models claim to be able to align object tokens with specific visual targets, but there are non-negligible gaps between the two.
Approach: They conduct diagnostic experiments to examine how the agents perceive multimodal input by ablation diagnostics input data.
Outcome: The results show that indoor and outdoor navigation agents refer to object and direction tokens when making decisions.
Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation (P19-1)

Copied to clipboard

Challenge: Existing metrics for vision-and-language navigation focus on goal completion rather than the sequence of actions corresponding to the instructions.
Approach: They propose to use a room-to-room dataset to measure the length of instruction followed by agents.
Outcome: The proposed metric outperforms existing metrics for Room-to-Room tasks because it is direct-to goal shortest.
LOViS: Learning Orientation and Visual Signals for Vision and Language Navigation (2022.coling-1)

Copied to clipboard

Challenge: Existing Transformer-based VLN agents entangle orientation and vision information, which limits the learning of each information source.
Approach: They propose to design a navigation agent with explicit Orientation and Vision modules . they use a set of pre-training tasks to feed the modules into the model .
Outcome: The proposed model improves on R2R and R4R datasets and achieves state-of-the-art results.
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents (2026.acl-long)

Copied to clipboard

Challenge: Vision-and-Language Navigation (VLN) is a subfield of embodied AI that integrates natural language understanding, visual perception, and sequential decision-making to allow autonomous agents to navigate and interact within visual environments.
Approach: They propose a modular framework that introduces structured, skill-based reasoning into Transformer-based VLN agents.
Outcome: The proposed framework decomposes navigation into atomic skills handled by a specialized agent.
Follow the Beaten Path: The Role of Route Patterns on Vision-Language Navigation Agents Generalization Abilities (2025.naacl-long)

Copied to clipboard

Challenge: Vision and language navigation (VLN) is a challenging task towards the creation of embodied agents.
Approach: They propose a solution that combines visual and linguistic features to enable VLN . they propose augmentation of the training data to fill the gap in missing patterns .
Outcome: The proposed solution fills the gap in missing patterns of training data.
Improving Cross-Modal Alignment in Vision Language Navigation via Syntactic Information (2021.naacl-main)

Copied to clipboard

Challenge: Existing vision language navigation tasks require soft attention over words to locate instructions . a new approach uses syntax information to ground instructions with visual information .
Approach: They propose a vision language navigation agent that utilizes syntax information to enhance alignment between the instruction and the current visual scenes.
Outcome: The proposed agent outperforms the baseline model that does not use syntax information on the Room-to-Room dataset, especially in the unseen environment.
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding (2020.emnlp-main)

Copied to clipboard

Challenge: Room-Across-Room (RxR) is a vision-and-language navigation dataset that addresses gaps in existing ones by addressing known biases in paths and eliciting more references to visible entities.
Approach: They introduce a new Vision-and-Language Navigation (VLN) dataset that addresses biases in paths and elicits more references to visible entities.
Outcome: The proposed model learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations.
Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions (2022.acl-long)

Copied to clipboard

Challenge: Vision-and-Language Navigation (VLN) is a research topic that is gaining attention in the field of artificial intelligence.
Approach: They propose to build an embodied agent that can communicate with humans in natural language and navigate in real 3D environments.
Outcome: This paper reviews current studies in the emerging field of vision-and-language navigation . it highlights limitations and opportunities for future work .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations