Challenge: Recent advances in large vision-language models have improved causal reasoning abilities . however, current models struggle with tasks like causal reasoning .
Approach: They propose a fine-grained and unified definition of causality involving interactions between humans and objects.
Outcome: The proposed model surpasses traditional commonsense causality by including explicit causal graphs . it also shows that current LVLMs can benefit from a causally inspired prompting strategy .

Similar Papers

CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large vision-language models have shown impressive ability in various language tasks, especially with their emergent in-context learning capability.
Approach: They propose a causal reasoning benchmark for multi-modal in-context learning from large vision-language models that incorporates visual inputs.
Outcome: The proposed model outperforms existing models on three visual causal reasoning tasks and demonstrates their strengths and weaknesses.
Large Language Models and Causal Inference in Collaboration: A Comprehensive Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown great potential to enhance Natural Language Processing (NLP) models in areas such as predictive accuracy, fairness, robustness, and explainability.
Approach: They evaluate or improve generative Large Language Models from a causal perspective in areas such as reasoning capacity, fairness and safety issues, explainability, and handling multimodality.
Outcome: The proposed models can be used to perform causal relationship discovery and causal effect estimation tasks.
Causal Inference with Large Language Model: A Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Existing causal inference frameworks do not match human judgment in several key areas, such as domain knowledge, logical inference, and cultural context.
Approach: They propose to apply large language models to causal inference tasks . they summarize the main causal problems and approaches and compare their results .
Outcome: The proposed methods are compared with traditional methods in healthcare, finance, and economics.
Can Large Language Models Infer Causal Relationships from Real-World Text? (2026.acl-long)

Copied to clipboard

Challenge: Existing work evaluating large language models relies on synthetic or simplified texts with explicit causal relationships.
Approach: They develop a benchmark to evaluate LLMs' ability to infer causal relationships from texts . they use a dataset of texts with different levels of explicitness and complexity .
Outcome: The proposed benchmark is the first-ever real-world dataset for this task.
Multimodal Causal Reasoning Benchmark: Challenging Multimodal Large Language Models to Discern Causal Links Across Modalities (2025.findings-acl)

Copied to clipboard

Challenge: Existing MLLMs lack robustness in multimodal causal reasoning compared to their performance in textual settings.
Approach: They propose a novel multimodal chain-of-thought (CoT) reasoning benchmark that leverages siamese images and text pairs to challenge MLLMs.
Outcome: The proposed benchmark leverages siamese images and text pairs to challenge MLLMs.
CausalGraph2LLM: Evaluating LLMs for Causal Queries (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have opened up new avenues for their use beyond standard Natural Language Processing tasks.
Approach: They propose a benchmark to evaluate the capabilities of Large Language Models (LLMs) they use over 700k queries to compare their encoding capabilities.
Outcome: The proposed benchmark compared LLMs on graph-level and node-level queries and open-sourced and closed models.
Causal-LLM: A Unified One-Shot Framework for Prompt- and Data-Driven Causal Graph Discovery (2025.findings-emnlp)

Copied to clipboard

Challenge: Current causal discovery methods rely on pairwise or iterative strategies that fail to capture global dependencies, amplify local biases, and reduce overall accuracy.
Approach: They propose a framework for one-step full causal graph discovery using prompt-based discovery and a data-driven method for settings without metadata.
Outcome: The proposed framework outperforms state-of-the-art models by approximately 40% in edge accuracy on datasets like Asia and Sachs while maintaining strong performance on more complex graphs.
CausalEval: Towards Better Causal Reasoning in Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been used for a variety of tasks, including problem-solving, decision-making, and understanding of the world.
Approach: They propose a review of existing methods aimed at enhancing LMs for causal reasoning . they categorize existing methods as reasoning engines or as helpers providing knowledge or data to traditional methods .
Outcome: The proposed methods perform better than existing methods on a range of tasks.
ACCESS : A Benchmark for Abstract Causal Event Discovery and Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for identifying event causality in NLP are limited in their scale and rely on lexical cues.
Approach: They propose a benchmark for identifying abstract causality from a large-scale dataset.
Outcome: The proposed benchmark can be leveraged for enhancing QA reasoning performance in LLMs.
CLEAR: Can Language Models Really Understand Causal Graphs? (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing language models lack a conceptual framework for understanding causal graphs, but there is still potential for improvement.
Approach: They develop a framework to define causal graph understanding by assessing language models’ behaviors through four practical criteria derived from diverse disciplines.
Outcome: The proposed framework defines three complexity levels and encompasses 20 causal graph-based tasks across 20 different levels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations