Challenge: Evidence suggests prelinguistic infants are capable of recognizing discrete events in real-world stimuli.
Approach: They propose a multimodal formulation for partially-defined events and cast the extraction of these events as a three-stage span retrieval task.
Outcome: The proposed approach can extract events from 14.5 hours of annotated current event videos and 1,168 text documents, containing 22.8K labeled event-centric entities.

Similar Papers

Multimedia Event Extraction with LLM Knowledge Editing (2025.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal event extraction methods focus on weakly aligning features from wellpretrained unimodal encoders, resulting in redundant feature perception.
Approach: They propose a multimodal event extraction strategy with a redundant feature selection mechanism that enhances event understanding ability of multimodal large language models.
Outcome: The proposed method outperforms the state-of-the-art (SOTA) baselines on the M2E2 benchmark.
Multi-Document Event Extraction Using Large and Small Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to multi-document event extraction have limited attention . despite its practical significance, this task has inherent challenges .
Approach: They propose a collaborative framework that integrates large language models for multi-step reasoning and fine-tuned small language models to handle key subtasks.
Outcome: The proposed framework outperforms existing methods and provides new insights into collaborative reasoning to tackle the complexities of multi-document event extraction.
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event Extraction (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for event extraction ignore motion representations in videos and are misguided by background noise.
Approach: They propose a text-video based multimodal event extraction framework that integrates video appearance features and motion representations with video appearance.
Outcome: The proposed framework outperforms the state-of-the-art methods in the event extraction field.
Multimodal Grounding for Language Processing (C18-1)

Copied to clipboard

Challenge: Recent developments in multimodal processing facilitate conceptual grounding of language.
Approach: They analyze multimodal processing to examine the benefits and challenges of multimodal grounding . they focus on multimodal linguistic grounding of verbs which play a crucial role in compositional power of language.
Outcome: The proposed methods improve the cognitive models of human information processing and address the challenges that arise.
Conundrums in Event Coreference Resolution: Making Sense of the State of the Art (2021.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen the successful application of span-based neural models to entity-based information extraction tasks such as entity coreference resolution (CR) Existing event coreference resolvers focused on feature engineering are few and far between, let alone event corefers.
Approach: They propose to adapt existing span-based event reference systems to event coreference by adapting the models originally developed for entity coreference to event CR.
Outcome: The proposed model improves the representations of entity mentions in entity-based IE tasks compared to non-span models .
Event Semantic Classification in Context (2024.findings-eacl)

Copied to clipboard

Challenge: In this work, we focus on the semantic classification of events in context to help machines gain a deeper understanding of events.
Approach: They propose to integrate event semantics into downstream tasks to help machines understand events better.
Outcome: The proposed model improves the understanding of events in context.
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks (2026.findings-acl)

Copied to clipboard

Challenge: Survey aims to identify challenges of multimodal unlearning for vision, language, audio and video . retraining after deletion requests or policy updates is often impractical, survey finds .
Approach: They propose to enable selective removal across modalities while retaining overall utility.
Outcome: This study compares models with existing models to identify weaknesses and improves performance.
Massively Multi-Lingual Event Understanding: Extraction, Visualization, and Search (2023.acl-demo)

Copied to clipboard

Challenge: Using only English training data, ISI-Clear makes global events available on-demand in 100 languages . Using a fixed task, events may still shift from day to day .
Approach: They propose a cross-lingual zero-shot event extraction system that makes global events available on-demand in 100 languages.
Outcome: The proposed system can extract events from non-English documents in 100 languages.
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities.
Approach: They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs.
Outcome: The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations