Challenge: In this paper, we evaluate use of different attribution methods for aiding identification of training data artifacts.
Approach: They propose hybrid methods that combine saliency maps and instance attribution methods to aid in identifying training data artifacts.
Outcome: The proposed methods can be used to efficiently uncover artifacts in training data when a challenging validation set is available.

Similar Papers

An Empirical Comparison of Instance Attribution Methods for NLP (2021.naacl-main)

Copied to clipboard

Challenge: Influence functions provide machinery for identifying training instances that may have led to a specific prediction, but are computationally expensive and prohibitive in many cases.
Approach: They evaluate the degree to which different potential instance attribution agrees with respect to the importance of training samples.
Outcome: The proposed methods exhibit desirable characteristics similar to more complex methods, but are computationally expensive.
Influence Tuning: Demoting Spurious Correlations via Instance Attribution and Instance-Driven Updates (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to interpret black-box models to learn spurious correlations are not well understood.
Approach: They propose a procedure that leverages model interpretations to update parameters towards a plausible interpretation rather than an interpretation that relies on spurious patterns in data.
Outcome: The proposed procedure outperforms baseline methods that use adversarial training in a controlled setup.
Mitigating Dataset Artifacts in Natural Language Inference Through Automatic Contextual Data Augmentation and Learning Optimization (2022.lrec-1)

Copied to clipboard

Challenge: In recent years, natural language inference has been an emerging research area . a new data augmentation technique is used to augment pre-trained language models .
Approach: They propose to combine automatic contextual data augmentation with a learning procedure for natural language inference.
Outcome: The proposed method outperforms baseline pre-trained language models on benchmark datasets and adversarial examples.
Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions (2020.acl-main)

Copied to clipboard

Challenge: Modern deep learning models for NLP are notoriously opaque, and this has motivated efforts to design example-specific approaches to interpret such models.
Approach: They propose to use influence functions to explain models by highlighting important words in input text to provide models with an explanation.
Outcome: The proposed approach is particularly useful for natural language inference, a task in which ‘saliency maps’ may not have clear interpretation.
Incorporating Priors with Feature Attribution on Text Classification (P19-1)

Copied to clipboard

Challenge: Feature attribution methods are used to help users interpret complex models.
Approach: They propose a feature attribution method that integrates feature attributed features into the objective function to allow machine learning practitioners to incorporate priors in model building.
Outcome: The proposed method reduces undesired model biases without a tradeoff on the original task and improves classifier performance in scarce data setting.
Locally Distributed Activation Vectors for Guided Feature Attribution (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to explain predictions of deep neural networks are unstable and do not always provide faithful explanations to the target model.
Approach: They propose a method to learn explanations-specific representations while constructing deep network models for text classification.
Outcome: The proposed method improves model interpretability while preserving predictive performance.
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods (2024.acl-long)

Copied to clipboard

Challenge: Language Models acquire parametric knowledge from their training process, embedding it within their weights.
Approach: They propose a new evaluation framework to quantify and compare the knowledge revealed by Instance Attribution and Neuron Attributions.
Outcome: The proposed evaluation framework compares the knowledge revealed by IA and NA with that of neuron attribution methods.
Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data (2022.acl-long)

Copied to clipboard

Challenge: Experimental results show that REtrieving from the traINing datA only can lead to significant gains on multiple NLG and NLU tasks.
Approach: They propose to retrieve training instances from traINing datA and concatenate them with input to generate output.
Outcome: The proposed method achieves state-of-the-art results on XSum, BigPatent, and CommonsenseQA.
Instance-Based Neural Dependency Parsing (2021.tacl-1)

Copied to clipboard

Challenge: Existing models that use instance-based inference for dependency parsing are difficult to understand for humans.
Approach: They develop neural models that adopt an interpretable inference process for dependency parsing.
Outcome: The proposed models achieve competitive accuracy with standard neural models and have plausibility of instance-based explanations.
Asymmetric feature interaction for interpreting model predictions (2023.findings-acl)

Copied to clipboard

Challenge: Prior work on feature interaction attribution studies focus on asymmetric interaction that only explains the additional influence of a set of words in combination, which fails to capture asymmetry influence that contributes to model prediction.
Approach: They propose an asymmetric feature interaction attribution explanation model that explores asymmetry higher-order feature interactions in the inference of deep neural NLP models.
Outcome: The proposed model outperforms state-of-the-art models on two sentiment classification datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations