Challenge: Current methods for instruction generation depend on privileged inputs such as semantic maps, landmark annotations, and panoramic views.
Approach: They propose a task that generates coherent navigation instructions from egocentric visual observations.
Outcome: The proposed task generates coherent navigation instructions from egocentric visual data . the proposed task improves performance over state-of-the-art methods in BLEU-4 and CIDEr scores .

Similar Papers

Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction (D18-1)

Copied to clipboard

Challenge: Existing models that map from inputs to actions are inefficient and require hand-crafted meaning representations.
Approach: They propose to decompose instruction execution to goal prediction and action generation . they introduce two benchmarks for instruction following: LANI and CHAI .
Outcome: The proposed model decomposes instruction execution to goal prediction and action generation.
Semantic Map-based Generation of Navigation Instructions (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to navigation instruction generation use a sequence of panorama images as visual input.
Approach: They propose a new approach to navigation instruction generation using semantic maps as visual input and frame it as an image captioning task.
Outcome: The proposed model is based on a dataset of a human vision and language navigation task and human subjects are asked to manually assess the quality of the generated instructions.
Can LLM’s Generate Human-Like Wayfinding Instructions? Towards Platform-Agnostic Embodied Instruction Synthesis (2024.naacl-short)

Copied to clipboard

Challenge: 83.3% of users find the synthesized instructions accurately capture the details of the environment and show characteristics similar to those of human-generated instructions.
Approach: They propose an algorithm that uses in-context learning to condition an LLM to generate instructions using just a few references.
Outcome: The proposed algorithm is platform-agnostic and 83.3% of users find it to be accurate and similar to human-generated instructions.
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization (2026.eacl-long)

Copied to clipboard

Challenge: Text2Vis systems generate functional code but resulting charts lack semantic alignment and clarity.
Approach: They propose a framework that integrates post-execution feedback with textual accuracy, code validity, and visualization quality.
Outcome: The proposed framework outperforms strong zero-shot and supervised baselines and shows robust generalization to out-of-domain datasets.
Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data (2024.findings-acl)

Copied to clipboard

Challenge: OpenAI's GPT-4 has demonstrated remarkable multimodal capabilities, but specific mechanics of GPT4 remain unknown.
Approach: They propose a data collection methodology that synchronously synthesizes images and dialogues for visual instruction tuning.
Outcome: The proposed method improves on ten commonly assessed models and provides greater flexibility compared to existing methods.
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, large multimodal models struggle with basic tasks like reading time from a clock and planning a route using a road map.
Approach: They propose a multimodal self-instruct that synthesizes massive abstract images and visual reasoning instructions.
Outcome: The proposed model synthesizes 11,193 abstract images and reasoning instructions across eight visual scenarios.
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM (2025.findings-acl)

Copied to clipboard

Challenge: High-performance vision-and-language navigation models require large amounts of training data, the high cost of manual annotating has seriously hindered this field.
Approach: They propose a retrieval-augmented generation framework that generates user demand instructions for vision-and-language navigation.
Outcome: The proposed model achieves SOTA performance on the REVERIE benchmark.
Visually-Guided Policy Optimization for Multimodal Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing RLVRs lack visual faithfulness due to text-dominated reasoning . a novel framework to reinforce visual focus during policy optimization is proposed .
Approach: They propose a framework to reinforce visual focus during policy optimization using visual attention compensation mechanism.
Outcome: The proposed framework exhibits better visual activation and superior performance in multimodal reasoning and visual-dependent tasks.
Show and Guide: Instructional-Plan Grounded Vision and Language Model (2024.emnlp-main)

Copied to clipboard

Challenge: Existing plans-following language models (LLMs) are not capable of multimodal input and output, resulting in inconsistent performance on multimodal tasks.
Approach: They propose a multimodal plan-following language model that integrates both textual plans and visual information to bring cross-modality to instructional tasks.
Outcome: The proposed model performs well on multimodal and textual dialogue in a plan-grounded setting.
Self-Guided Alignment: Adaptive Preference Sensing for Multi-Objective Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to align LLMs with diverse human values rely on ground-truth scores . existing approaches implicitly approximate an average-user preference, thereby failing to capture heterogeneity of human values or accommodate conflicting user needs.
Approach: They propose a framework that transforms passive reward dependency into an intrinsic adaptive sensing capability.
Outcome: The proposed framework outperforms state-of-the-art models in multiple model scales and improves preference alignment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations