Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing explanation methods for image classification struggle to provide faithful and plausible explanations for predictions. |
| Approach: | They propose a natural language explanation method that can be applied to any CNN-based classifier without altering its training process or affecting predictive performance. |
| Outcome: | The proposed method can be applied to any CNN-based classifier without altering its training process or affecting predictive performance. |
Similar Papers
Pipeline for modeling causal beliefs from natural language (2023.acl-demo)
Copied to clipboard
| Challenge: | Existing methods to analyze language data for psychological causality are difficult to advance as they do not isolate cognitive mechanisms. |
| Approach: | They propose a pipeline that leverages a Large Language Model to identify causal claims made in natural language documents and applies a clustering algorithm to group causal claims based on their semantic topics. |
| Outcome: | The proposed pipeline analyzes the Covid-19 vaccine in tweets and generates a causal claim network. |
Learning to Faithfully Rationalize by Construction (2020.acl-main)
Copied to clipboard
| Challenge: | Neural models dominate NLP but it remains difficult to know why they make specific predictions for sequential text inputs. |
| Approach: | They propose a model to produce faithful rationales for neural text classification by defining independent snippet extraction and prediction modules. |
| Outcome: | The proposed model produces faithful explanations even when the model is complex and complex. |
Interpreting Recurrent and Attention-Based Neural Models: a Case Study on Natural Language Inference (D18-1)
Copied to clipboard
| Challenge: | In this paper, we examine the behavior of deep learning models in their intermediate layers . saliency determines what is critical for the final decision of a deep model . |
| Approach: | They propose to interpret the intermediate layers of deep models by visualizing the saliency of attention and LSTM gating signals. |
| Outcome: | The proposed methods reveal interesting insights and identify critical information contributing to the model decisions. |
Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work has focused on generating convincing rationales that appear to be subjectively faithful, but it remains unclear whether these explanations are epistemic faithful. |
| Approach: | They propose a method that enhances epistemic faithfulness by guiding explanation generation through attention-level interventions, informed by token-level heatmaps. |
| Outcome: | The proposed method significantly improves epistemic faithfulness across multiple models, benchmarks, and prompts. |
On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods for explaining "black-box" models such as Influence Functions are becoming more popular. |
| Approach: | They propose a semantic-based evaluation metric that can better align with humans’ judgment of explanations than the widely adopted diagnostic or re-training measures. |
| Outcome: | The proposed method can better align with humans’ judgment of explanations than diagnostic or re-training measures. |
NILE : Natural Language Inference with Faithful Natural Language Explanations (2020.acl-main)
Copied to clipboard
| Challenge: | Recent growth in popularity of deep learning models on NLP classification tasks has accompanied the need for generating some form of natural language explanation of predicted labels. |
| Approach: | They propose a novel method which generates labels along with its faithful explanations. |
| Outcome: | The proposed method is more accurate than previously reported methods and has higher sensitivity than previous methods. |
Systematic Analysis of Image Schemas in Natural Language through Explainable Multilingual Neural Language Processing (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for automatic detection of image schemas in natural language rely on specific assumptions about word classes as indicators of spatio-temporal events. |
| Approach: | They propose to train a supervised classifier that classifies natural language expressions into image schemas using a large dataset of examples from image schema literature. |
| Outcome: | The proposed model performs best in German, Russian, and French, and is based on a small dataset of examples from image schema literature. |
Human-grounded Evaluations of Explanation Methods for Text Classification (D19-1)
Copied to clipboard
| Challenge: | Explainable Artificial Intelligence (XAI) is aimed at providing explanations for decisions made by AI systems. |
| Approach: | They propose to use model-agnostic and model-specific explanation methods for CNNs for text classification to provide human-grounded evaluations. |
| Outcome: | The proposed methods could be used to explain models' results and improve AIs and humans in many cases. |
Why Attention is Not Explanation: Surgical Intervention and Causal Reasoning about Neural Models (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study finds brittleness in explanations obtained through attention mechanisms . a philosophy of science theory allows robust yet non-causal reasoning in explanation . |
| Approach: | They propose to use philosophy of science to examine the state-of-the-art in explanation for NLP models . they argue that it is impossible to explain attention-based learning by attention mechanisms . |
| Outcome: | The proposed model selection criteria are based on philosophy of science theories . the proposed model is based upon a model that is more explainable than a classical model . |
An Empirical Study on Explanations in Out-of-Domain Settings (2022.acl-long)
Copied to clipboard
| Challenge: | Recent work in Natural Language Processing has focused on extracting faithful explanations . yet, little is known about how post-hoc explanations perform in out-of-domain settings . |
| Approach: | They propose to use a random baseline to evaluate out-of-domain post-hoc explanation faithfulness . they suggest select-then-predict models demonstrate comparable predictive performance in out- of-domain settings to full-text trained models. |
| Outcome: | The proposed models perform better in out-of-domain settings than full-text models. |