MM-IGLU: Multi-Modal Interactive Grounded Language Understanding (2024.lrec-main)
Copied to clipboard
| Challenge: | In human-robot interaction, a robot interprets user commands related to its environment, aiming to discern whether a specific command can be executed. |
| Approach: | They propose to integrate user statements with environment's description to create a multi-modal interactive Grounded language understanding model that integrates both visual and textual data. |
| Outcome: | The proposed model integrates user’s statement with environment’s description and a cutting-edge Multi-Modal Large Language Model merges both visual and textual data. |
Similar Papers
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback (2024.findings-eacl)
Copied to clipboard
| Challenge: | In many approaches to Natural Language Processing tasks, language is inherently interactive. |
| Approach: | They propose to use human-AI collaboration to improve human-human interaction by providing feedback that the agent can understand and utilize. |
| Outcome: | The proposed task is an interactive grounded language understanding task in a MineCraft-like world. |
Training Multi-Modal LLMs through Dialogue Planning for HRI (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to enhance Multi-Modal Large Language Models (MLLMs) with explicit dialogue planning improves response accuracy and quality, and allows models trained in one language to transfer effectively to another. |
| Approach: | They propose an approach that enhances Multi-Modal Large Language Models with a novel explicit dialogue planning phase that allows agents to refine their understanding of ambiguous commands. |
| Outcome: | The proposed approach reduces hallucinations and improves task feasibility by fine-tuning and assessing Multi-Modal models in human-robot interaction scenarios. |
Grounded Semantic Role Labelling from Synthetic Multimodal Data for Situated Robot Commands (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing symbolic parsers lack flexibility to operate in complex, dynamic environments. |
| Approach: | They propose a framework that combines frame semantics with perceptual grounding to enable robots to interpret commands via multimodal logical forms. |
| Outcome: | The proposed framework produces over 11,000 image-command pairs and lowers the cost of manual parsers. |
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)
Copied to clipboard
Pei Fu, Tongkun Guan, Zining Wang, Zhentao Guo, Chen Duan, Hao Sun, Boming Chen, Qianyi Jiang, Jiayao Ma, Kai Zhou, Junfeng Luo
| Challenge: | Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms. |
| Approach: | They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks. |
| Outcome: | The proposed models perform well on mainstream benchmarks and are compared with other models. |
GroundingGPT: Language Enhanced Multi-modal Grounding Model (2024.acl-long)
Copied to clipboard
Zhaowei Li, Qi Xu, Dong Zhang, Hang Song, YiQing Cai, Qi Qi, Ran Zhou, Junting Pan, Zefeng Li, Vu Tu, Zhida Huang, Tao Wang
| Challenge: | Existing multi-modal large language models focus on capturing global information while neglecting the fine-grained local information in multimodal inputs. |
| Approach: | They propose an end-to-end language enhanced multi-modal grounding model that performs fine-grained grounding tasks for image, video and audio. |
| Outcome: | The proposed model achieves impressive fine-grained understanding of multi-modal inputs while maintaining or improving its global comprehension capabilities. |
MM-LLMs: Recent Advances in MultiModal Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | MultiModal Large Language Models (MM-LLMs) have undergone significant advances in the past year . traditional MM models incur substantial computational costs, especially when trained from scratch . |
| Approach: | They propose a taxonomy encompassing 126 MM-LLMs and summarize key training recipes to enhance their potency. |
| Outcome: | The proposed models preserve the reasoning and decision-making capabilities of LLMs and empower diverse range of MM tasks. |
Explainability and Interpretability of Multilingual Large Language Models: A Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing literature on multilingual large language models lacks transparency in their internal processes. |
| Approach: | They propose to use multilingual large language models to examine their explainability and interpretability methods. |
| Outcome: | The present study examines the explainability and interpretability of multilingual large language models. |
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Prior studies show that large language models map multilingual content into English-aligned representations at intermediate layers before projecting them back into target-language token spaces in the later layers. |
| Approach: | They propose a method to identify and manipulate dimensions that are sparse and sparsity-based . they propose to use as few as 50 sentences of either parallel or monolingual data to manipulate these dimensions . |
| Outcome: | Experiments on a multilingual generation control task show the interpretability of these dimensions. |
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives (2024.findings-acl)
Copied to clipboard
Thong Nguyen, Yi Bin, Junbin Xiao, Leigang Qu, Yicong Li, Jay Zhangjie Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
| Challenge: | Existing video-language understanding systems with human-like senses can mimic both our linguistic medium and visual environment with temporal dynamics. |
| Approach: | They propose to develop video-language understanding systems with human-like senses . they summarize their methods and highlight challenges associated with them . |
| Outcome: | The proposed models perform well in a variety of tasks and domains. |
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies. |
| Approach: | They propose a psychometric evaluation pipeline grounded in realistic human-AI interactions to probe value orientations and novel tasks for evaluating value understanding in an open-ended value space. |
| Outcome: | The proposed evaluation pipeline is grounded in realistic human-AI interactions and performs tasks that approximate expert conclusions in value-related extraction and generation tasks. |