Challenge: In human-robot interaction, a robot interprets user commands related to its environment, aiming to discern whether a specific command can be executed.
Approach: They propose to integrate user statements with environment's description to create a multi-modal interactive Grounded language understanding model that integrates both visual and textual data.
Outcome: The proposed model integrates user’s statement with environment’s description and a cutting-edge Multi-Modal Large Language Model merges both visual and textual data.

Similar Papers

Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback (2024.findings-eacl)

Copied to clipboard

Challenge: In many approaches to Natural Language Processing tasks, language is inherently interactive.
Approach: They propose to use human-AI collaboration to improve human-human interaction by providing feedback that the agent can understand and utilize.
Outcome: The proposed task is an interactive grounded language understanding task in a MineCraft-like world.
Training Multi-Modal LLMs through Dialogue Planning for HRI (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to enhance Multi-Modal Large Language Models (MLLMs) with explicit dialogue planning improves response accuracy and quality, and allows models trained in one language to transfer effectively to another.
Approach: They propose an approach that enhances Multi-Modal Large Language Models with a novel explicit dialogue planning phase that allows agents to refine their understanding of ambiguous commands.
Outcome: The proposed approach reduces hallucinations and improves task feasibility by fine-tuning and assessing Multi-Modal models in human-robot interaction scenarios.
Grounded Semantic Role Labelling from Synthetic Multimodal Data for Situated Robot Commands (2025.emnlp-main)

Copied to clipboard

Challenge: Existing symbolic parsers lack flexibility to operate in complex, dynamic environments.
Approach: They propose a framework that combines frame semantics with perceptual grounding to enable robots to interpret commands via multimodal logical forms.
Outcome: The proposed framework produces over 11,000 image-command pairs and lowers the cost of manual parsers.
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms.
Approach: They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks.
Outcome: The proposed models perform well on mainstream benchmarks and are compared with other models.
GroundingGPT: Language Enhanced Multi-modal Grounding Model (2024.acl-long)

Copied to clipboard

Challenge: Existing multi-modal large language models focus on capturing global information while neglecting the fine-grained local information in multimodal inputs.
Approach: They propose an end-to-end language enhanced multi-modal grounding model that performs fine-grained grounding tasks for image, video and audio.
Outcome: The proposed model achieves impressive fine-grained understanding of multi-modal inputs while maintaining or improving its global comprehension capabilities.
MM-LLMs: Recent Advances in MultiModal Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: MultiModal Large Language Models (MM-LLMs) have undergone significant advances in the past year . traditional MM models incur substantial computational costs, especially when trained from scratch .
Approach: They propose a taxonomy encompassing 126 MM-LLMs and summarize key training recipes to enhance their potency.
Outcome: The proposed models preserve the reasoning and decision-making capabilities of LLMs and empower diverse range of MM tasks.
Explainability and Interpretability of Multilingual Large Language Models: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Existing literature on multilingual large language models lacks transparency in their internal processes.
Approach: They propose to use multilingual large language models to examine their explainability and interpretability methods.
Outcome: The present study examines the explainability and interpretability of multilingual large language models.
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior studies show that large language models map multilingual content into English-aligned representations at intermediate layers before projecting them back into target-language token spaces in the later layers.
Approach: They propose a method to identify and manipulate dimensions that are sparse and sparsity-based . they propose to use as few as 50 sentences of either parallel or monolingual data to manipulate these dimensions .
Outcome: Experiments on a multilingual generation control task show the interpretability of these dimensions.
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives (2024.findings-acl)

Copied to clipboard

Challenge: Existing video-language understanding systems with human-like senses can mimic both our linguistic medium and visual environment with temporal dynamics.
Approach: They propose to develop video-language understanding systems with human-like senses . they summarize their methods and highlight challenges associated with them .
Outcome: The proposed models perform well in a variety of tasks and domains.
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies.
Approach: They propose a psychometric evaluation pipeline grounded in realistic human-AI interactions to probe value orientations and novel tasks for evaluating value understanding in an open-ended value space.
Outcome: The proposed evaluation pipeline is grounded in realistic human-AI interactions and performs tasks that approximate expert conclusions in value-related extraction and generation tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations