Challenge: Existing models that understand narratives should infer these implicit states and their causal relationships with the narrative's explicit events.
Approach: They propose a dataset that contains inferable participant states, a counterfactual perturbation to each state and the changes to the story that would be necessary if the counterfact was true.
Outcome: The proposed model can reason about the impact of changes to the story that would be necessary if the counterfactual were true.

Similar Papers

Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives? (2026.eacl-long)

Copied to clipboard

Challenge: Contemporary models of (human) reading comprehension characterize comprehension as a dynamic process in which the reader continually builds and updates representations to maintain coherence and integrate new information with prior knowledge.
Approach: They use a paired narrative dataset to examine the extent to which large language models can reliably separate incoherent and coherent stories.
Outcome: The proposed models do not eliminate the deficits in the model internal state and behavior.
Counterfactual Story Reasoning and Generation (D19-1)

Copied to clipboard

Challenge: a desired property of AI systems is counterfactual reasoning: ability to predict causal changes in future events.
Approach: They propose to rewrite a short story and a counterfactual event to make it compatible with the given counterfact.
Outcome: The proposed task requires deep understanding of causal narrative chains and counterfactual invariance . the proposed dataset includes 81,407 counterfact "branches" without a rewritten storyline .
SOCCER: An Information-Sparse Discourse State Tracking Collection in the Sports Commentary Domain (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for state tracking are limited and state changes are less densely distributed over utterances.
Approach: They propose to turn to simplified, fully observable systems that show some of these properties.
Outcome: The proposed system shows that state changes occur infrequently while messages are "chatter" it allows for rich descriptions of state while avoiding the complexities of other settings.
A Causal Approach for Counterfactual Reasoning in Narratives (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for counterfactual reasoning in narratives are based on dataset-specific heuristics, but they are abusing unique patterns, i.e., the feature of minimum editing, in the dataset, which limits the generality of their methods.
Approach: They propose a basic VAE module for counterfactual reasoning in narratives and introduce a pre-trained classifier and external event commonsense to mitigate the posterior collapse problem.
Outcome: The proposed method improves the causality between the counterfactual condition and the generated counterf actual outcome on two public benchmarks.
EvEntS ReaLM: Event Reasoning of Entity States via Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to model event implications fail to reason about the world, despite their knowledge of physical attributes.
Approach: They propose to use a model prompting technique to prompt models of event implications by targeting their understanding of physical attributes.
Outcome: The proposed model prompting technique is especially useful for unseen attributes or when only limited data is available.
TellMeWhy: A Dataset for Answering Why-Questions in Narratives (2021.findings-acl)

Copied to clipboard

Challenge: Existing models do not have the ability to answer "why" questions that require commonsense knowledge external to the narrative.
Approach: They propose a crowd-sourced dataset that asks why characters perform actions . they show that state-of-the-art models are far below human performance on answering such questions .
Outcome: The proposed dataset shows that state-of-the-art models are far below human performance on answering such questions.
CRASS: A Novel Data Set and Benchmark to Test Counterfactual Reasoning of Large Language Models (2022.lrec-1)

Copied to clipboard

Challenge: CRASS data set and benchmark provide novel test scheme to evaluate large language models . authors present and explain the CRAS data set, a novel basis to test reasoning and natural language understanding of LLMs .
Approach: They introduce a new test scheme utilizing questionized counterfactual conditionals to evaluate large language models.
Outcome: The proposed model sets out to be the most powerful and valid tool to evaluate large language models.
A Method for Building a Commonsense Inference Dataset based on Basic Events (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to acquire commonsense are limited by the general-purpose language models.
Approach: They propose a method for building a commonsense inference dataset using crowdsourcing and automatic extraction from a corpus.
Outcome: The proposed method can solve 104k commonsense inference problems in a Japanese corpus with high accuracy, but low bias.
Understanding the Language of Political Agreement and Disagreement in Legislative Texts (2020.acl-main)

Copied to clipboard

Challenge: Despite the fact that state-level legislation is rarely discussed, it has a dramatic influence on the everyday life of residents of the respective states.
Approach: They propose a large-scale dataset linking state bills and legislator information, geographical information about their districts, and donations and donors’ information.
Outcome: The proposed model improves over strong text-based models by integrating the state-level text and the legislative context.
BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on the use of Large Language Models (LLMs) focus primarily on English, leaving the cross-lingual generalization of aligned behavior underexplored.
Approach: They propose a structured generator-extractor pipeline and a multi-dimensional distributional analysis framework to examine how narrative attributes vary across languages, models, and social conditions.
Outcome: The proposed model reveals substantial cross-lingual variability in narrative generation patterns, indicating that distributions observed in English do not always exhibit similar characteristics in other languages, particularly in lower-resource settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations