Challenge: Existing models for translating free-form natural language instructions to a high-level plan for behavioral robot navigation are difficult due to the variability in the way people describe routes.
Approach: They propose an end-to-end deep learning model for translating free-form natural language instructions to a high-level plan for robot navigation.
Outcome: The proposed model significantly outperforms baseline approaches on a new dataset containing 10,050 pairs of navigation instructions.

Similar Papers

Sub-Instruction Aware Vision-and-Language Navigation (2020.emnlp-main)

Copied to clipboard

Challenge: Despite significant advances, few previous works are able to fully utilize the strong correspondence between visual and textual sequences.
Approach: They propose to provide agents with fine-grained annotations during training and provide them with sub-instructions and their corresponding paths.
Outcome: The proposed method improves the performance of four state-of-the-art agents in a room-to-room (R2R) benchmark dataset.
LLM-driven Instruction Following: Progresses and Concerns (2023.emnlp-tutorial)

Copied to clipboard

Challenge: a tutorial on task instruction is aimed at researchers and practitioners interested in NLP generalization . labeled examples are unlikely to be available in large numbers or do not exist .
Approach: This tutorial will examine the progress of natural language processing (NLP) using labeled examples. authors propose that task instructions act as a novel resource for supervision.
Outcome: This tutorial aims to answer questions about instruction-driven NLP . it focuses on the use of task instructions in a low-shot scenario .
VLN-Trans: Translator for the Vision and Language Navigation Agent (2023.acl-long)

Copied to clipboard

Challenge: We observe two kinds of instructions that make the grounding in the vision-and-language navigation task quite challenging.
Approach: They propose to use a translator module to convert instructions into easy-to-follow sub-instruction representations at each step.
Outcome: The proposed model is based on a Room2Room (R2R), Room4room (R4R), and Room2room Last (R1R-Last) datasets and achieves state-of-the-art results on multiple benchmarks.
Visually-Grounded Planning without Vision: Language Models Infer Detailed Plans from High-level Instructions (2020.findings-emnlp)

Copied to clipboard

Challenge: Currently, the best performing systems can complete less than 1% of tasks successfully . we show that it is possible to generate gold multi-step plans from language directives alone without any visual input in 26% of unseen cases .
Approach: They propose a task that requires a virtual robot to translate natural language directives into detailed sequences of actions that accomplish a goal in a real environment.
Outcome: The proposed task generates gold multi-step plans from natural language directives in 26% of unseen cases.
Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)

Copied to clipboard

Challenge: Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation.
Approach: This tutorial introduces the techniques used in cutting-edge research on speech translation.
Outcome: The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages.
X-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual Instructions (2024.findings-acl)

Copied to clipboard

Challenge: Large language models respond well in high-resource languages but struggle in low-resourced languages.
Approach: They propose a method to construct cross-lingual instruction following samples with instruction in English and response in low-resource languages.
Outcome: The proposed method builds a large-scale cross-lingual instruction tuning dataset on 10 languages.
NavA3: Understanding Any Instruction, Navigating Anywhere, Finding Anything (2026.acl-long)

Copied to clipboard

Challenge: Existing embodied navigation methods struggle with such tasks due to their limitations in comprehending high-level human instructions and localizing objects with an open vocabulary.
Approach: They propose a hierarchical framework for long-horizon navigation that integrates human instructions with 3D scene views.
Outcome: The proposed model achieves SOTA results and can complete long-horizon navigation tasks across different robot embodiments in real-world environments.
GesNavi: Gesture-guided Outdoor Vision-and-Language Navigation (2024.eacl-srw)

Copied to clipboard

Challenge: Existing datasets for outdoor Vision-and-Language Navigation (VLN) tasks do not include verbal instructions for communicating with mobility.
Approach: They propose a dataset for gesture-guided outdoor VLN instructions with demonstrative expressions that incorporates gestures and linguistic commands.
Outcome: The proposed datasets are compared against existing datasets and analysed in detail.
Generating Landmark Navigation Instructions from Maps as a Graph-to-Text Problem (2021.acl-long)

Copied to clipboard

Challenge: Current navigation services based on turns and distances of named streets . humans use efficient mode of navigation based around visible and salient landmarks .
Approach: They propose a neural model that takes OpenStreetMap representations as input and learns to generate navigation instructions that contain salient landmarks from human natural language instructions.
Outcome: The proposed model can generate navigation instructions that contain salient landmarks from openStreetMap . it is based on a dataset of 7,672 instances verified by human navigation in Street View .
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout (N19-1)

Copied to clipboard

Challenge: Existing approaches perform significantly worse in unseen environments compared to seen ones.
Approach: They propose to use a ‘environmental dropout’ method to generate unseen triplets to generate new paths and instructions to generalize the agent.
Outcome: The proposed agent outperforms the state-of-the-art approaches on the private unseen test set and is ranked top on the leaderboard.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations