Challenge: Several models have been proposed to interpret models with feature-based interpretability methods.
Approach: They propose to quantify the robustness of neural text classifiers by using two randomization tests to compare models with identical initializations.
Outcome: The proposed methods show surprising deviations from expected behavior . the results raise questions about the extent of insights that practitioners may draw from interpretations.

Similar Papers

Interpretability and Analysis in Neural NLP (2020.acl-tutorials)

Copied to clipboard

Challenge: a tutorial aims to introduce the nascent field of interpretability and analysis of neural networks in NLP .
Approach: This tutorial will introduce the nascent field of interpretability and analysis of neural networks in NLP.
Outcome: This tutorial will introduce the nascent field of interpretability and analysis of neural networks in NLP.
Interpreting the Robustness of Neural NLP Models to Textual Perturbations (2022.findings-acl)

Copied to clipboard

Challenge: Modern Natural Language Processing models are sensitive to input perturbations and their performance can decrease when applied to noisy data.
Approach: They propose to explain the extent to which a model is affected by an unseen textual perturbation by the learnability of the perturbation.
Outcome: The proposed model is better at identifying a perturbation (higher learnability) but worse at ignoring it (lower robustness).
Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the Transformer (2022.emnlp-main)

Copied to clipboard

Challenge: Neural Machine Translation (NMT) relies on source sentence and target prefix attributions for each input token.
Approach: They propose an interpretability method that tracks input tokens’ attributions for both contexts and extends it to any encoder-decoder Transformer-based model.
Outcome: The proposed method can be extended to any encoder-decoder Transformer-based model and provides insights into their behaviour.
Transformer-specific Interpretability (2024.eacl-tutorials)

Copied to clipboard

Challenge: Transformers are dominant play-ers in various scientific fields, but their inner workings remain opaque.
Approach: This tutorial presents a trending approach to interpreting Transformers . it uses specific features of the Transformer architecture to quantify context- mixing interactions .
Outcome: This tutorial aims to show how a new trending approach can be applied to Transformer-based models.
Demystifying Neural Fake News via Linguistic Feature-Based Interpretation (2022.coling-1)

Copied to clipboard

Challenge: Recent advances to neural fake news generators have made it difficult to understand how misinformation generated by these models may best be confronted.
Approach: They conduct feature-based analysis to gain an interpretative understanding of the linguistic attributes that neural fake news generators may most effectively exploit.
Outcome: The proposed models are compared with models trained on subsets of features and confronted with increasingly advanced neural fake news.
A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies show that fine-tuned textual transformer models are vulnerable to adversarial text perturbations.
Approach: They extract 13 different features representing a wide range of input fine-tuning corpora properties and use them to predict adversarial robustness of the fine- tuned models.
Outcome: The proposed framework can be used as an additional tool for robustness evaluation since it saves 30x-193x runtime compared to the traditional technique and can be easily used under adversarial training.
Robust Text Classification: Analyzing Prototype-Based Networks (2024.findings-emnlp)

Copied to clipboard

Challenge: Language models exhibit a drop in performance on noisy data, which can cause classifiers to incorrectly change their predictions.
Approach: They propose to use Prototype-Based Networks to classify examples based on their similarity to prototypical examples of a class (prototypes) they show that PBNs offer more robustness under both targeted and static adversarial attacks.
Outcome: The proposed model is robust to noise and targets both targeted and static attacks.
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders (2026.findings-eacl)

Copied to clipboard

Challenge: Concept-based explanations for large language models are not well understood in text classification.
Approach: They propose a model with a specialized classifier head and activation rate sparsity loss for sentence classification . they compare it to existing models with HI-Concept and ConceptShap .
Outcome: The proposed model improves both the causality and interpretability of the extracted features.
Evaluating Saliency Methods for Neural Language Models (2021.naacl-main)

Copied to clipboard

Challenge: a general complaint of neural network models is that their internal decision mechanisms are hard to understand.
Approach: They evaluate the quality of prediction interpretations from two perspectives: plausibility and faithfulness.
Outcome: The evaluation of saliency methods on neural language models shows they can be trusted . the methods can be used to interpret the same prediction, but they disagree on interpretations .
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)

Copied to clipboard

Challenge: Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies.
Approach: They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances .
Outcome: The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations