Challenge: Existing explanation methods for image classification struggle to provide faithful and plausible explanations for predictions.
Approach: They propose a natural language explanation method that can be applied to any CNN-based classifier without altering its training process or affecting predictive performance.
Outcome: The proposed method can be applied to any CNN-based classifier without altering its training process or affecting predictive performance.

Similar Papers

Pipeline for modeling causal beliefs from natural language (2023.acl-demo)

Copied to clipboard

Challenge: Existing methods to analyze language data for psychological causality are difficult to advance as they do not isolate cognitive mechanisms.
Approach: They propose a pipeline that leverages a Large Language Model to identify causal claims made in natural language documents and applies a clustering algorithm to group causal claims based on their semantic topics.
Outcome: The proposed pipeline analyzes the Covid-19 vaccine in tweets and generates a causal claim network.
Learning to Faithfully Rationalize by Construction (2020.acl-main)

Copied to clipboard

Challenge: Neural models dominate NLP but it remains difficult to know why they make specific predictions for sequential text inputs.
Approach: They propose a model to produce faithful rationales for neural text classification by defining independent snippet extraction and prediction modules.
Outcome: The proposed model produces faithful explanations even when the model is complex and complex.
Interpreting Recurrent and Attention-Based Neural Models: a Case Study on Natural Language Inference (D18-1)

Copied to clipboard

Challenge: In this paper, we examine the behavior of deep learning models in their intermediate layers . saliency determines what is critical for the final decision of a deep model .
Approach: They propose to interpret the intermediate layers of deep models by visualizing the saliency of attention and LSTM gating signals.
Outcome: The proposed methods reveal interesting insights and identify critical information contributing to the model decisions.
Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance (2026.acl-long)

Copied to clipboard

Challenge: Prior work has focused on generating convincing rationales that appear to be subjectively faithful, but it remains unclear whether these explanations are epistemic faithful.
Approach: They propose a method that enhances epistemic faithfulness by guiding explanation generation through attention-level interventions, informed by token-level heatmaps.
Outcome: The proposed method significantly improves epistemic faithfulness across multiple models, benchmarks, and prompts.
On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for explaining "black-box" models such as Influence Functions are becoming more popular.
Approach: They propose a semantic-based evaluation metric that can better align with humans’ judgment of explanations than the widely adopted diagnostic or re-training measures.
Outcome: The proposed method can better align with humans’ judgment of explanations than diagnostic or re-training measures.
NILE : Natural Language Inference with Faithful Natural Language Explanations (2020.acl-main)

Copied to clipboard

Challenge: Recent growth in popularity of deep learning models on NLP classification tasks has accompanied the need for generating some form of natural language explanation of predicted labels.
Approach: They propose a novel method which generates labels along with its faithful explanations.
Outcome: The proposed method is more accurate than previously reported methods and has higher sensitivity than previous methods.
Systematic Analysis of Image Schemas in Natural Language through Explainable Multilingual Neural Language Processing (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for automatic detection of image schemas in natural language rely on specific assumptions about word classes as indicators of spatio-temporal events.
Approach: They propose to train a supervised classifier that classifies natural language expressions into image schemas using a large dataset of examples from image schema literature.
Outcome: The proposed model performs best in German, Russian, and French, and is based on a small dataset of examples from image schema literature.
Human-grounded Evaluations of Explanation Methods for Text Classification (D19-1)

Copied to clipboard

Challenge: Explainable Artificial Intelligence (XAI) is aimed at providing explanations for decisions made by AI systems.
Approach: They propose to use model-agnostic and model-specific explanation methods for CNNs for text classification to provide human-grounded evaluations.
Outcome: The proposed methods could be used to explain models' results and improve AIs and humans in many cases.
Why Attention is Not Explanation: Surgical Intervention and Causal Reasoning about Neural Models (2020.lrec-1)

Copied to clipboard

Challenge: a recent study finds brittleness in explanations obtained through attention mechanisms . a philosophy of science theory allows robust yet non-causal reasoning in explanation .
Approach: They propose to use philosophy of science to examine the state-of-the-art in explanation for NLP models . they argue that it is impossible to explain attention-based learning by attention mechanisms .
Outcome: The proposed model selection criteria are based on philosophy of science theories . the proposed model is based upon a model that is more explainable than a classical model .
An Empirical Study on Explanations in Out-of-Domain Settings (2022.acl-long)

Copied to clipboard

Challenge: Recent work in Natural Language Processing has focused on extracting faithful explanations . yet, little is known about how post-hoc explanations perform in out-of-domain settings .
Approach: They propose to use a random baseline to evaluate out-of-domain post-hoc explanation faithfulness . they suggest select-then-predict models demonstrate comparable predictive performance in out- of-domain settings to full-text trained models.
Outcome: The proposed models perform better in out-of-domain settings than full-text models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations