Papers by Björn Buchhold
XAI-Attack: Utilizing Explainable AI to Find Incorrectly Learned Patterns for Black-Box Adversarial Example Creation (2024.lrec-main)
Copied to clipboard
| Challenge: | Adversarial examples can be used to trick machine learning models into making erroneous predictions, causing poorer insights and lower confidence in the information gathered. |
| Approach: | They propose a textual adversarial example method that identifies falsely learned word indicators by leveraging explainable AI methods as importance functions on incorrectly predicted instances. |
| Outcome: | The proposed method outperforms existing examples and training methods and shows baseline improvements of up to 23 percentage points on adversarial tasks. |