Challenge: a number of studies have reported high accuracies in NLP tasks due to simple heuristics and dataset artifacts.
Approach: They use a case where two words/terms occur in a shared context to construct a causal diagram . they also investigate the robustness to irrelevant changes and sensitivity to impactful changes of Transformers .
Outcome: The proposed method bolsters the fact that similar benchmark accuracy scores may be observed for models that exhibit very different behaviour.

Similar Papers

Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond (2022.tacl-1)

Copied to clipboard

Challenge: causality has not had the same importance in natural language processing, says aaron e. smith . he says research on causality in NLP remains scattered across domains without unified definitions .
Approach: They propose to consolidate research on causality in NLP across academic areas . they explore potential uses of causal inference to improve robustness, fairness, interpretability .
Outcome: The proposed method is a unified overview of causal inference for the NLP community.
Developmental Negation Processing in Transformer Language Models (2022.acl-short)

Copied to clipboard

Challenge: Negation is an important construct in language for reasoning over the truth of propositions, garnering interest from philosophy (Horn, 1989) and psycholinguistics (Zwaan, 2012).
Approach: They propose to frame a natural language inference task as a problem and examine how well transformers can process negation categories.
Outcome: The proposed models perform better on certain categories, suggesting clear differences in how they are processed.
CausalNLP Tutorial: An Introduction to Causality for Natural Language Processing (2022.emnlp-tutorials)

Copied to clipboard

Challenge: Establishing causal relationships is a fundamental goal of scientific research . lack of clear definitions, notations, benchmark datasets, and challenges remains .
Approach: They introduce the fundamentals of causal discovery and causal effect estimation to the natural language processing audience and provide an overview of causal perspectives to NLP problems.
Outcome: This tutorial introduces the fundamentals of causal discovery and causal effect estimation to the natural language processing audience and provides an overview of causal perspectives to NLP problems.
Causal Effects of Linguistic Properties (2021.naacl-main)

Copied to clipboard

Challenge: Social scientists have long been interested in the causal effects of language, studying questions like: How should political candidates describe their personal history to appeal to voters?
Approach: They propose an algorithm for estimating causal effects of linguistic properties that leverages distant supervision and a pre-trained language model to adjust for the text.
Outcome: The proposed method outperforms other methods when estimating the effect of Amazon review sentiment on semi-simulated sales figures.
Causal Inference with Large Language Model: A Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Existing causal inference frameworks do not match human judgment in several key areas, such as domain knowledge, logical inference, and cultural context.
Approach: They propose to apply large language models to causal inference tasks . they summarize the main causal problems and approaches and compare their results .
Outcome: The proposed methods are compared with traditional methods in healthcare, finance, and economics.
Large Language Models and Causal Inference in Collaboration: A Comprehensive Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown great potential to enhance Natural Language Processing (NLP) models in areas such as predictive accuracy, fairness, robustness, and explainability.
Approach: They evaluate or improve generative Large Language Models from a causal perspective in areas such as reasoning capacity, fairness and safety issues, explainability, and handling multimodality.
Outcome: The proposed models can be used to perform causal relationship discovery and causal effect estimation tasks.
Simple but Challenging: Natural Language Inference Models Fail on Simple Sentences (2022.findings-emnlp)

Copied to clipboard

Challenge: Natural language inference (NLI) tasks are difficult to perform on large datasets . a small number of simple sentences can improve model performance, authors say .
Approach: They propose to use syntactically simple sentences to test the inference ability of NLI models.
Outcome: The proposed set of simple sentences shows that the models fine-tuned on MNLI and SNLI perform poorly on Simple Pair.
CausalEval: Towards Better Causal Reasoning in Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been used for a variety of tasks, including problem-solving, decision-making, and understanding of the world.
Approach: They propose a review of existing methods aimed at enhancing LMs for causal reasoning . they categorize existing methods as reasoning engines or as helpers providing knowledge or data to traditional methods .
Outcome: The proposed methods perform better than existing methods on a range of tasks.
Investigating the Performance of Transformer-Based NLI Models on Presuppositional Inferences (2022.coling-1)

Copied to clipboard

Challenge: Presuppositions are assumptions that are taken for granted by an utterance.
Approach: They propose to use heuristics to create alternative "contrastive" test cases . they also analyze samples from ImpPres datasets to better understand their predictions .
Outcome: The proposed model performs better on the ImpPres dataset than on the other datasets.
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models (2024.eacl-long)

Copied to clipboard

Challenge: Recent studies have indicated that NLI models have an understanding of lexical and compositional semantics.
Approach: They propose a framework to assess the extent of semantic sensitivity in NLI models . they use adversarially generated examples with minor semantics-preserving surface-form variations .
Outcome: The proposed framework shows that NLI models struggle with minor variations requiring knowledge of compositional semantics .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations