| Challenge: | a tutorial aims to build agents that understand language using a simulated environment . situated reasoning is a critical aspect of human language understanding . |
| Approach: | This tutorial combines a synthesis of multimodal grounding and meaning representation techniques with formal and computational models of situated reasoning. |
| Outcome: | This tutorial combines multimodal grounding and meaning representation techniques with formal and computational models of embodied reasoning. |
Similar Papers
Pragmatics in Language Grounding: Phenomena, Tasks, and Modeling Approaches (2023.findings-emnlp)
Copied to clipboard
| Challenge: | People rely heavily on context to enrich meaning beyond what is literally said. |
| Approach: | They analyze how task goals, environmental contexts, and communicative affordances in each work enrich linguistic meaning. |
| Outcome: | The proposed frameworks are based on linguistic goals, environmental contexts, and communicative affordances to enrich linguistic meaning. |
Multimodal Grounding for Language Processing (C18-1)
Copied to clipboard
| Challenge: | Recent developments in multimodal processing facilitate conceptual grounding of language. |
| Approach: | They analyze multimodal processing to examine the benefits and challenges of multimodal grounding . they focus on multimodal linguistic grounding of verbs which play a crucial role in compositional power of language. |
| Outcome: | The proposed methods improve the cognitive models of human information processing and address the challenges that arise. |
Grounding Language in Multi-Perspective Referential Communication (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using a dataset of 2,970 human-written referring expressions, we find that the performance of automated models in both reference generation and comprehension lags behind that of pairs of human agents. |
| Approach: | They propose a task and dataset for referring expression generation and comprehension in multi-agent embodied environments where two agents must take into account one another's visual perspective to produce and understand references to objects in a scene. |
| Outcome: | The proposed model outperforms the strongest proprietary model and improves communicative success from 58.9 to 69.3% when trained with a listener. |
Learning Language through Grounding (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | This tutorial provides a historical overview of grounding and discusses its use in computational linguistics and in computational language processing. |
| Approach: | They introduce the concept of grounding and discuss future directions and open challenges . they will delve into recent progress in learning lexical semantics, syntax, and complex meanings through various forms of ground. |
| Outcome: | This course will provide an overview of the field of grounding and discuss future directions and challenges related to large language models and scaling. |
Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models. |
| Approach: | They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. |
| Outcome: | The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. |
Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction (2021.naacl-main)
Copied to clipboard
| Challenge: | Scholarly work in this area uses toy worlds and synthetic linguistic data, but grounded language learning offers several practical and scientific advantages. |
| Approach: | They propose to model teacher-learner dynamics through natural interactions occurring between users and search engines. |
| Outcome: | The proposed model is better than non-grounded models on compositionality and zero-shot inference tasks. |
Incorporating Visual Semantics into Sentence Representations within a Grounded Space (D19-1)
Copied to clipboard
| Challenge: | Language grounding is an active field aiming at enriching textual representations with visual information. |
| Approach: | They propose to transfer visual information to textual representations by learning an intermediate representation space: the grounded space. |
| Outcome: | The proposed model outperforms the previous state-of-the-art on classification and semantic relatedness tasks. |
Visual Grounding Helps Learn Word Meanings in Low-Data Regimes (2024.naacl-long)
Copied to clipboard
| Challenge: | Modern neural language models (LMs) require distinctly un-human-like ways to achieve these results. |
| Approach: | They train a diverse set of LM architectures with and without auxiliary visual supervision on datasets of varying scales. |
| Outcome: | The proposed models exhibit better learning of syntactic categories, lexical relations, semantic features, word similarity and alignment with human neural representations. |
Language (Re)modelling: Towards Embodied Language Understanding (2020.acl-main)
Copied to clipboard
| Challenge: | Despite the rapid progress in NLU, current systems lack the rich mental representations that people use for language understanding. |
| Approach: | They propose an approach to representation and learning based on the tenets of embodied cognitive linguistics (ECL) they propose a system architecture along with a roadmap towards realizing this vision. |
| Outcome: | The proposed approach will improve the performance of existing systems and provide a roadmap towards realizing this vision. |
A Formal Analysis of Multimodal Referring Strategies Under Common Ground (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on multimodality in the CL/NLP community, but it has not been widely studied. |
| Approach: | They propose to analyze mixed-modality definite referring expressions using gestures and linguistic descriptions. |
| Outcome: | The proposed models can predict viewer judgment of referring expressions and generate more natural and informative expressions. |