Multimodal Contextualized Semantic Parsing from Speech (2024.acl-long)

Copied to clipboard

Challenge: Towards this goal, we introduce Semantic Parsing in Contextual Environments (SPICE) task designed to enhance artificial agents’ contextual awareness by integrating multimodal inputs with prior contexts.
Approach: They introduce a task designed to enhance artificial agents’ contextual awareness by integrating multimodal inputs with prior contexts.
Outcome: The proposed task is based on the VG-SPICE dataset and the Audio-Vision Dialogue Scene Parser (AViD-SP) it allows agents to maintain their contextual state within a structured, dense information framework that is scalable and interpretable .

Similar Papers

Context Dependent Semantic Parsing: A Survey (2020.coling-main)

Copied to clipboard

Challenge: Semantic parsing is the task of translating natural language utterances into machine-readable meaning representations.
Approach: They propose to use contextual information to translate natural language utterances into machine-readable meaning representations.
Outcome: The proposed methods do not utilize contextual information, which could boost the semantic parsing systems.
Semantic Parsing for Conversational Question Answering over Knowledge Graphs (2023.eacl-main)

Copied to clipboard

Challenge: Recent years have seen an increasing number of applications aiming to build conversational interfaces based on information retrieval and user recommendation.
Approach: They develop a dataset where user questions are annotated with Sparql parses and system answers correspond to execution results thereof.
Outcome: The proposed parsers can be used to ground questions into queries over definitions in a knowledge graph with large vocabularies.
Data Augmentation with Paraphrase Generation and Entity Extraction for Multimodal Dialogue System (2022.lrec-1)

Copied to clipboard

Challenge: Contextually aware intelligent agents are often required to understand the users and their surroundings in real-time.
Approach: They propose to build a multimodal dialogue system for children learning basic math concepts using limited datasets.
Outcome: The proposed system improves the Natural Language Understanding (NLU) module of a task-oriented SDS pipeline with limited dataset resources.
Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on explicit context, but do not address context-dependent pragmatic understanding.
Approach: They propose a benchmark for evaluating ISA understanding through integrated reasoning over visual context and dialogue.
Outcome: Experiments show that state-of-the-art models struggle with visually grounded indirect speech acts . linguistic meaning emerges through the relationship between an utterance and situational context .
Situated and Interactive Multimodal Conversations (2020.coling-main)

Copied to clipboard

Challenge: Situated Interactive MultiModal Conversations (SIMMC) is a new direction for virtual assistants that handle multimodal inputs and perform multimodal actions.
Approach: They propose to use Situated Interactive MultiModal Conversations (SIMMC) to train agents to take multimodal actions grounded in a co-evolving multimodal context.
Outcome: The proposed model will be made publicly available.
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models fail to adapt to unfamiliar speakers and language varieties . however, there are significant gaps in the adaptation of certain varieties based on the test speaker, variety, or recording conditions .
Approach: They propose a framework that allows for in-context learning in Phi-4 Multimodal . they find that as few as 12 example utterances reduce word error rates by 19.7% .
Outcome: The proposed framework reduces word error rates by 19.7% across diverse English corpora.
Multimodal Grounding for Language Processing (C18-1)

Copied to clipboard

Challenge: Recent developments in multimodal processing facilitate conceptual grounding of language.
Approach: They analyze multimodal processing to examine the benefits and challenges of multimodal grounding . they focus on multimodal linguistic grounding of verbs which play a crucial role in compositional power of language.
Outcome: The proposed methods improve the cognitive models of human information processing and address the challenges that arise.
Contextual Semantic Parsing for Multilingual Task-Oriented Dialogues (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods for predicting state of a conversation are limited to a few languages . a method that can be applied to other languages will benefit the large population of speakers of many other languages.
Approach: They propose to automatically translate large-scale dialogue data sets in one language to produce an effective semantic parser for other languages using machine translation.
Outcome: The proposed model reduces the compounding effect of translation errors without harming the accuracy in practice.
Coarse-to-Fine Decoding for Neural Semantic Parsing (P18-1)

Copied to clipboard

Challenge: Experimental results show that semantic parsing is more efficient than using simple decoders.
Approach: They propose a structure-aware neural architecture which decomposes the semantic parsing process into two stages.
Outcome: The proposed architecture consistently improves performance on four datasets characteristic of different domains and meaning representations.
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for semantics discovery focus on text, video, and audio, failing to leverage the rich multimodal information in the real world.
Approach: They propose a method to construct augmentation views for multimodal data and use them to perform pre-training to establish well-initialized representations for subsequent clustering.
Outcome: The proposed method improves on benchmark multimodal intent and dialogue act datasets by 2-6% over state-of-the-art methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations