Papers by Jens Tuyls
AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models (D19-3)
Copied to clipboard
| Challenge: | Existing interpretation codebases make it difficult to apply these methods to new models and tasks. |
| Approach: | They propose a framework for interpreting NLP models that provides explanations for specific models. |
| Outcome: | The proposed framework provides interpretation primitives for any AllenNLP model and task, a suite of built-in interpretation methods, and a library of front-end visualization components. |
Gradient-based Analysis of NLP Models is Manipulable (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work has shown that explanation techniques can be unstable and can be manipulated to hide the actual reasoning behind the predictions of NLP models. |
| Approach: | They propose to merge a BERT-based sentiment classifier with a Facade Model that overwhelms the gradients without affecting the predictions. |
| Outcome: | The proposed model overwhelms the gradients without affecting the predictions on a variety of NLP tasks, such as sentiment analysis, NLI, and QA. |