GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning (2026.findings-acl)
Copied to clipboard
Fengyi Wu, Yifei Dong, Yilong Dai, Guangyu Chen, Qifeng Wu, Huiting Huang, Hang Wang, Qi Dai, Alexander G Hauptmann, Zhi-Qi Cheng
| Challenge: | Current methods for instruction generation depend on privileged inputs such as semantic maps, landmark annotations, and panoramic views. |
| Approach: | They propose a task that generates coherent navigation instructions from egocentric visual observations. |
| Outcome: | The proposed task generates coherent navigation instructions from egocentric visual data . the proposed task improves performance over state-of-the-art methods in BLEU-4 and CIDEr scores . |
Similar Papers
Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction (D18-1)
Copied to clipboard
| Challenge: | Existing models that map from inputs to actions are inefficient and require hand-crafted meaning representations. |
| Approach: | They propose to decompose instruction execution to goal prediction and action generation . they introduce two benchmarks for instruction following: LANI and CHAI . |
| Outcome: | The proposed model decomposes instruction execution to goal prediction and action generation. |
Semantic Map-based Generation of Navigation Instructions (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to navigation instruction generation use a sequence of panorama images as visual input. |
| Approach: | They propose a new approach to navigation instruction generation using semantic maps as visual input and frame it as an image captioning task. |
| Outcome: | The proposed model is based on a dataset of a human vision and language navigation task and human subjects are asked to manually assess the quality of the generated instructions. |
Can LLM’s Generate Human-Like Wayfinding Instructions? Towards Platform-Agnostic Embodied Instruction Synthesis (2024.naacl-short)
Copied to clipboard
| Challenge: | 83.3% of users find the synthesized instructions accurately capture the details of the environment and show characteristics similar to those of human-generated instructions. |
| Approach: | They propose an algorithm that uses in-context learning to condition an LLM to generate instructions using just a few references. |
| Outcome: | The proposed algorithm is platform-agnostic and 83.3% of users find it to be accurate and similar to human-generated instructions. |
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization (2026.eacl-long)
Copied to clipboard
| Challenge: | Text2Vis systems generate functional code but resulting charts lack semantic alignment and clarity. |
| Approach: | They propose a framework that integrates post-execution feedback with textual accuracy, code validity, and visualization quality. |
| Outcome: | The proposed framework outperforms strong zero-shot and supervised baselines and shows robust generalization to out-of-domain datasets. |
Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data (2024.findings-acl)
Copied to clipboard
Yanda Li, Chi Zhang, Gang Yu, Wanqi Yang, Zhibin Wang, Bin Fu, Guosheng Lin, Chunhua Shen, Ling Chen, Yunchao Wei
| Challenge: | OpenAI's GPT-4 has demonstrated remarkable multimodal capabilities, but specific mechanics of GPT4 remain unknown. |
| Approach: | They propose a data collection methodology that synchronously synthesizes images and dialogues for visual instruction tuning. |
| Outcome: | The proposed method improves on ten commonly assessed models and provides greater flexibility compared to existing methods. |
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model (2024.emnlp-main)
Copied to clipboard
Wenqi Zhang, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu, Yueting Zhuang
| Challenge: | Using large language models, large multimodal models struggle with basic tasks like reading time from a clock and planning a route using a road map. |
| Approach: | They propose a multimodal self-instruct that synthesizes massive abstract images and visual reasoning instructions. |
| Outcome: | The proposed model synthesizes 11,193 abstract images and reasoning instructions across eight visual scenarios. |
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM (2025.findings-acl)
Copied to clipboard
| Challenge: | High-performance vision-and-language navigation models require large amounts of training data, the high cost of manual annotating has seriously hindered this field. |
| Approach: | They propose a retrieval-augmented generation framework that generates user demand instructions for vision-and-language navigation. |
| Outcome: | The proposed model achieves SOTA performance on the REVERIE benchmark. |
Visually-Guided Policy Optimization for Multimodal Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing RLVRs lack visual faithfulness due to text-dominated reasoning . a novel framework to reinforce visual focus during policy optimization is proposed . |
| Approach: | They propose a framework to reinforce visual focus during policy optimization using visual attention compensation mechanism. |
| Outcome: | The proposed framework exhibits better visual activation and superior performance in multimodal reasoning and visual-dependent tasks. |
Show and Guide: Instructional-Plan Grounded Vision and Language Model (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing plans-following language models (LLMs) are not capable of multimodal input and output, resulting in inconsistent performance on multimodal tasks. |
| Approach: | They propose a multimodal plan-following language model that integrates both textual plans and visual information to bring cross-modality to instructional tasks. |
| Outcome: | The proposed model performs well on multimodal and textual dialogue in a plan-grounded setting. |
Self-Guided Alignment: Adaptive Preference Sensing for Multi-Objective Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to align LLMs with diverse human values rely on ground-truth scores . existing approaches implicitly approximate an average-user preference, thereby failing to capture heterogeneity of human values or accommodate conflicting user needs. |
| Approach: | They propose a framework that transforms passive reward dependency into an intrinsic adaptive sensing capability. |
| Outcome: | The proposed framework outperforms state-of-the-art models in multiple model scales and improves preference alignment. |