Challenge: Existing models of event processing do not understand the essentiality of step events towards a goal event.
Approach: They propose to deconstruct a goal event into a discrete representation of finer-grained (step) events, which are not equally important to the goal.
Outcome: The proposed model can understand the essentiality of different step events towards a goal event.

Similar Papers

Combining Tradition with Modernness: Exploring Event Representations in Vision-and-Language Models for Visual Goal-Step Inference (2023.acl-srw)

Copied to clipboard

Challenge: Existing methods for representing procedural knowledge are limited to capturing the most crucial information, namely actions and the participants, to learn stereotypical event sequences.
Approach: They propose a task that uses images to identify steps towards achieving a goal in the multimodal domain.
Outcome: The proposed task uses images that represent steps towards achieving a textually expressed goal in the multimodal domain.
The Imperfective Paradox in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing models rely on surface-level probabilistic heuristics to grasp compositional semantics of events . authors: current open-weight models operate as predictive narrative engines rather than faithful reasoners .
Approach: They propose a diagnostic dataset to probe the imperfective paradox . they uncover a pervasive Teleological Bias in open-weight models .
Outcome: The proposed dataset reveals a pervasive Teleological Bias in open-weight models . the findings suggest that these models operate as predictive narrative engines rather than faithful reasoners .
Reasoning about Goals, Steps, and Temporal Ordering with WikiHow (2020.emnlp-main)

Copied to clipboard

Challenge: Existing datasets focus on relation between procedural events, but little attention has been paid to relation between events.
Approach: They propose a set of reasoning tasks targeting goal-step relations and step-step temporal relations based on wikiHow articles . their automatically-generated training set allows models to transfer to out-of-domain tasks requiring knowledge of procedural events .
Outcome: The proposed dataset improves on SWAG, Snips, and Story Cloze Test in zero- and few-shot settings.
Event-Centric Natural Language Processing (2021.acl-tutorials)

Copied to clipboard

Challenge: This tutorial will provide an introduction to various methods for automating the extraction, conceptualization and prediction of events and their relations.
Approach: This tutorial will provide an introduction to various methods for automating events and their relations, and a wide range of NLU and commonsense understanding tasks.
Outcome: This tutorial will provide an introduction to various methods for automating extraction, conceptualization and prediction of events and their relations, and a wide range of NLU and commonsense understanding tasks.
Improving Event Definition Following For Zero-Shot Event Detection (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches on zero-shot event detection train models on datasets annotated with known event types and prompt them with unseen event definitions.
Approach: They propose to train models to better follow event definitions by using an automatic generated Diverse Event Definition dataset.
Outcome: The proposed model outperforms existing models on three open benchmarks on zero-shot event detection.
Conundrums in Event Coreference Resolution: Making Sense of the State of the Art (2021.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen the successful application of span-based neural models to entity-based information extraction tasks such as entity coreference resolution (CR) Existing event coreference resolvers focused on feature engineering are few and far between, let alone event corefers.
Approach: They propose to adapt existing span-based event reference systems to event coreference by adapting the models originally developed for entity coreference to event CR.
Outcome: The proposed model improves the representations of entity mentions in entity-based IE tasks compared to non-span models .
PizzaCommonSense: A Dataset for Commonsense Reasoning about Intermediate Steps in Cooking Recipes (2024.findings-emnlp)

Copied to clipboard

Challenge: Understanding procedural texts is essential for enabling machines to follow instructions and reason about tasks.
Approach: They propose a corpus of cooking recipes enriched with descriptions of intermediate steps . they propose enabling machines to follow instructions and reason about tasks .
Outcome: The proposed model achieves only 26% human-evaluated preference for generations . pizzaCommonsense is a benchmark for the reasoning capabilities of large language models .
Debiasing Event Understanding for Visual Commonsense Tasks (2022.findings-acl)

Copied to clipboard

Challenge: a recent study shows that object-based event understanding is purely likelihood-based, leading to incorrect event prediction.
Approach: They propose to mitigate object-based event understanding by optimizing aggregation with association-based prediction.
Outcome: The proposed approach improves visual commonsense reasoning tasks by combining do-calculus with association-based prediction.
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators (2026.findings-acl)

Copied to clipboard

Challenge: State-of-the-art NLP models are expensive and inefficient for event annotation.
Approach: They propose to integrate LLMs into a holistic workflow that summarizes news with event coreference resolution and argument extraction in three modes: AI-only, AI assistance, and human only.
Outcome: The proposed workflow integrates LLMs to alleviate human labor in a holistic pipeline.
Evaluating Step-by-step Reasoning Traces: A Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation practices are inconsistent, resulting in fragmented progress across evaluator design and benchmark development.
Approach: a survey provides a comprehensive overview of step-by-step reasoning evaluation . existing evaluation practices are inconsistent, resulting in fragmented progress .
Outcome: The proposed evaluation criteria are based on four top-level categories . the results are presented in a systematic review of the literature.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations