Challenge: Existing studies on language models' ability to explain their decisions in natural language have focused on self-generated counterfactual explanations (SCEs).
Approach: They evaluate whether LLMs can generate valid counterfactuals and minimal ones . authors suggest that SCEs are, at best, an ineffective explainability tool .
Outcome: The proposed language models can explain their decisions in natural language, the study finds . the models can produce valid counterfactual explanations, but make small edits that fail to change predictions.

Similar Papers

Can LLMs Explain Themselves Counterfactually? (2025.emnlp-main)

Copied to clipboard

Challenge: Explanations are an important tool for gaining insights into model behavior, calibrating user trust, and ensuring compliance.
Approach: They propose to use self-explanation to prompt models to explain outputs . they find that LLMs struggle to generate SCEs - their prediction often does not agree with their own counterfactual reasoning.
Outcome: The proposed methods can generate SCEs across families, sizes, temperatures, and datasets.
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable performance in NLP tasks, but their efficacy in generating high-quality CFs remains uncertain.
Approach: They compare LLMs' ability to generate CFs that flip the original label and human CF's.
Outcome: The proposed models generate fluent CFs, but struggle to keep the induced changes minimal.
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation (2026.acl-long)

Copied to clipboard

Challenge: Large language models excel at generating English counterfactuals but their effectiveness in generating multilingual counterfacts remains unclear.
Approach: They conduct automatic evaluations on both directly generated and derived counterfactuals in six languages and find that cross-lingual perturbations follow common strategic principles.
Outcome: The proposed models show that translation-based counterfactuals offer higher validity than their directly generated counterparts, but still fall short of matching the quality of the original English counterf actuals.
Self-Critique and Refinement for Faithful Natural Language Explanations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work has demonstrated that Large Language Models (LLMs) can self-critique and refine their initial outputs, but this capability remains unexplored for improving explanation faithfulness.
Approach: They propose a framework that enables models to improve the faithfulness of their own explanations through an iterative critique and refinement process without external supervision.
Outcome: The proposed framework reduces unfaithfulness rates in three datasets and four state-of-the-art LLMs by 36% compared to 54.81% for baseline.
Reliability of Distribution Predictions by LLMs: Insights from Counterintuitive Pseudo-Distributions (2025.naacl-srw)

Copied to clipboard

Challenge: Recent studies highlight the use of Large Language Models (LLMs) for predicting response distributions as a cost-effective survey method.
Approach: They examine whether LLMs can rationally estimate distributions when presented with explanations that are against commonsense.
Outcome: The proposed models can rationally estimate distributions when presented with explanations that are against commonsense, but smaller or less human-optimized models follow explanations uncritically, compared to larger models that resist counterintuitive explanations by leveraging their pretraining-acquired knowledge.
Prompting Large Language Models for Counterfactual Generation: An Empirical Study (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have made remarkable progress in a wide range of natural language understanding and generation tasks, but their ability to generate counterfactuals has not been examined systematically.
Approach: They propose a framework to evaluate LLMs' ability to generate counterfactuals based on key factors including intrinsic properties and prompt design.
Outcome: The proposed framework examines the strengths and weaknesses of large language models (LLMs) and identifies factors that influence their ability to generate counterfactuals.
Are self-explanations from Large Language Models faithful? (2024.findings-acl)

Copied to clipboard

Challenge: Instruction-tuned Large Language Models excel at many tasks and will explain their reasoning, so-called self-explanations.
Approach: They propose to employ self-consistency checks to measure faithfulness to LLMs to determine if they are model-dependent and if their reasoning is convincing and wrong.
Outcome: The proposed measures show that self-explanations are explanation, model, and task-dependent and should not be trusted in general.
Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations (2024.emnlp-main)

Copied to clipboard

Challenge: Autoregressive Large Language Models (LLMs) have demonstrated "emergent abilities" such as in-context learning, instruction following and reasoning.
Approach: They propose a method that generates rationales from post hoc explanation methods applied to small language models to improve their own performance.
Outcome: The proposed method improves on four SLMs and five datasets with strong reasoning abilities.
Faithfulness Tests for Natural Language Explanations (2023.acl-short)

Copied to clipboard

Challenge: Existing methods for explaining neural models are misleading as they often present reasons that are unfaithful to the model’s inner workings.
Approach: They propose a counterfactual input editor for inserting reasons that lead to counterfacts but are not reflected by the NLEs.
Outcome: The proposed model can evaluate emerging NLE models, proving a fundamental tool in the development of faithful explanations.
Counterfactuals of Counterfactuals: a back-translation-inspired approach to analyse counterfactual editors (2023.findings-acl)

Copied to clipboard

Challenge: Existing explanations for classifiers are counterfactual or contrastive . lack of universal ground truth for counterf actual edits hinders their evaluation .
Approach: They propose a back translation-inspired evaluation methodology that utilises earlier outputs of the explainer as ground truth proxies to investigate the consistency of explainers.
Outcome: The proposed method can provide valuable insights into the behaviour of predictor and explainer models and infer patterns that would otherwise be obscured.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations