HeadlineCause: A Dataset of News Headlines for Detecting Causalities (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets focus on commonsense causal reasoning or explicit causal relations . authors present dataset for detecting implicit causal relations between news headlines .
Approach: They present a dataset for detecting implicit causal relations between news headlines . they use 5000 headline pairs from English news and 9000 from Russian news .
Outcome: The proposed dataset shows that it is valid and can be used to predict implicit causal relations between headline pairs.

Similar Papers

The Causal News Corpus: Annotating Causal Relations in Event Sentences from News (2022.lrec-1)

Copied to clipboard

Challenge: Existing annotation guidelines for event causality focus on only explicit relations or clauses.
Approach: They propose an annotation schema for event causality that addresses these concerns . they annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not.
Outcome: The proposed annotation schema for event causality addresses these concerns . it performs well with 81.20% F1 score on test set and 83.46% in 5-folds cross-validation .
ClimateCause: Complex and Implicit Causal Structures in Climate Reports (2026.findings-acl)

Copied to clipboard

Challenge: Existing datasets for causal discovery from text lack granularity and abstraction for domains characterized by such complex causality.
Approach: They propose a manually expert-annotated dataset of higher-order causal structures from science-for-policy climate reports.
Outcome: The proposed dataset is highly readable and can be used to quantify readability.
Modeling Document-level Causal Structures for Event Causal Relation Identification (N19-1)

Copied to clipboard

Challenge: a study aims to identify all the event causal relations in a document, both within a sentence and across sentences . main challenges for achieving comprehensive causal relation identification are sparse among all possible event pairs . few causal relations are explicitly stated, especially for identifying cross-sentence causal relations .
Approach: They propose to identify all event causal relations in a document, both within a sentence and across sentences.
Outcome: The proposed model improves the performance of causal relation identification . it shows that the model can be used to identify cross-sentence causal relations .
A Review of Dataset and Labeling Methods for Causality Extraction (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for causal relationship extraction are limited and lack of unified methods hinder progress in the field.
Approach: They propose to summarize existing methods and propose a new causal sequence label method . they propose to use multiple candidate causal label sequences according to label controversy .
Outcome: The proposed method summarises existing methods and explores their practicability and extensibility from multiple perspectives.
A Case Study on Neural Headline Generation for Editing Support (N19-2)

Copied to clipboard

Challenge: a news-aggregator is a website or mobile application that aggregates web content . dozens of professional editors manually create their headlines, which are much shorter than the original headlines.
Approach: They propose a neural headline generation model that automatically generates short headlines from news articles.
Outcome: The proposed model is deployed to an editing support tool and compares editors' behavior before and after the release.
A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse (2025.emnlp-main)

Copied to clipboard

Challenge: Existing datasets focus on explicit causality in structured text, providing limited support for detecting implicit causal expressions.
Approach: They propose a dataset of Reddit posts annotated across four causal tasks . they use a binary causal classification, explicit vs. implicit causality, cause–effect span extraction and causal gist generation to bridge causal detection and reasoning over informal discourse.
Outcome: The proposed dataset analyzes 10,120 Reddit posts discussing public health related to the COVID-19 pandemic.
Improving Truthfulness of Headline Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing studies on abstractive summarization report ROUGE scores, but are concerned about the truthfulness of generated summaries.
Approach: They propose to remove untruthful instances from supervision data to improve headline generation . they build a binary classifier that predicts an entailment relation between an article and its headline .
Outcome: The proposed model improves on two popular datasets.
Identifying Predictive Causal Factors from News Streams (D19-1)

Copied to clipboard

Challenge: Existing word embedding techniques are not suited to learn relationships between words in different documents and contexts.
Approach: They propose a new framework to uncover the relationship between news events and real world phenomena by measuring how word occurrence influences future occurrence.
Outcome: The proposed framework outperforms existing methods in stock price prediction errors for 12 months and 4 years.
Can Large Language Models Infer Causal Relationships from Real-World Text? (2026.acl-long)

Copied to clipboard

Challenge: Existing work evaluating large language models relies on synthetic or simplified texts with explicit causal relationships.
Approach: They develop a benchmark to evaluate LLMs' ability to infer causal relationships from texts . they use a dataset of texts with different levels of explicitness and complexity .
Outcome: The proposed benchmark is the first-ever real-world dataset for this task.
Misinfo Reaction Frames: Reasoning about Readers’ Reactions to News Headlines (2022.acl-long)

Copied to clipboard

Challenge: Empirical results confirm that it is indeed possible for neural models to predict the prominent patterns of readers’ reactions to previously unseen news headlines.
Approach: They propose a pragmatic formalism for modeling how readers might react to a news headline . they propose 'misinfo' frames, which can be used to model reader perceptions of news reliability .
Outcome: The proposed model can predict readers' reactions to previously unseen headlines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations