Papers by Sanjiv Das
On the Lack of Robust Interpretability of Neural Text Classifiers (2021.findings-acl)
Copied to clipboard
Muhammad Bilal Zafar, Michele Donini, Dylan Slack, Cedric Archambeau, Sanjiv Das, Krishnaram Kenthapadi
| Challenge: | Several models have been proposed to interpret models with feature-based interpretability methods. |
| Approach: | They propose to quantify the robustness of neural text classifiers by using two randomization tests to compare models with identical initializations. |
| Outcome: | The proposed methods show surprising deviations from expected behavior . the results raise questions about the extent of insights that practitioners may draw from interpretations. |