Challenge: Existing work on explaining classifier decisions has not addressed local feature redundancy . a common way to explain why a model classified an example is to extract a sparse subset of features that were responsible for the decision .
Approach: They propose an adversarial method for producing high-recall explanations of text classifier decisions . they use a method which scans the residual of attention for remaining predictive signal .
Outcome: The proposed method produces high-recall explanations of text classifier decisions . it uses a set of human-annotated personal attacks to evaluate the impact .

Similar Papers

Computational Ad Hominem Detection (P19-2)

Copied to clipboard

Challenge: ad hominem attacks are introduced in debates as an easy win, but their impact on argumentation is limited . a machine learning approach to detect the personal attack is insufficient, we show .
Approach: They propose a machine learning approach that detects ad hominem attacks using social media data . they propose TF-IDF approaches that are insufficient to detect the personal attack .
Outcome: The proposed method has a recall of 80% for a social media data source.
On the Robustness of Self-Attentive Models (P19-1)

Copied to clipboard

Challenge: Experimental results show that self-attentive neural models are more robust against adversarial perturbations compared to recurrent neural networks.
Approach: They propose an adversarial attack algorithm that generates more natural adversarials . they propose to use the attention mechanism to learn a context-dependent representation .
Outcome: The proposed attack algorithm generates more natural adversarial examples that could mislead models but not humans.
Emotion Detection with Neural Personal Discrimination (D19-1)

Copied to clipboard

Challenge: Existing approaches to automatically predict the emotions of posts consider each post individually and predict their emotions independently.
Approach: They propose a Neural Personal Discrimination approach to identify personal attributes from posts and connect relevant posts with similar attributes to jointly learn their emotions.
Outcome: The proposed approach improves on existing models by capturing attributes-aware words and predicting emotions among relevant posts.
Mining Tweets that refer to TV programs with Deep Neural Networks (D19-55)

Copied to clipboard

Challenge: opinion mining is a popular natural language processing technique, but a problem is robustness for user-generated texts . a recent study shows that a model that handles context can extract the opinion target with 90% accuracy .
Approach: They propose a model that handles context in many natural language processing areas to solve a problem of extracting opinion references from text.
Outcome: Experiments on tweets that refer to television programs show the proposed model can extract opinion references with more than 90% accuracy.
Evaluating and Enhancing the Robustness of Neural Network-based Dependency Parsing Models with Adversarial Examples (2020.acl-main)

Copied to clipboard

Challenge: Previously studies focused on semantic tasks such as sentiment analysis, question answering and reading comprehension.
Approach: They propose two approaches to study where and how adversarial examples exist in dependency parsing . they use a state-of-the-art parser to find adversarials in existing texts .
Outcome: The proposed approaches show that adversarial examples exist in dependency parsing . they show that up to 77% of input examples admit adversarials .
Self-regulation: Employing a Generative Adversarial Network to Improve Event Detection (P18-1)

Copied to clipboard

Challenge: Recent studies show that neural networks can be used for event detection but can be contaminated by spurious features.
Approach: They propose a self-regulated learning approach by utilizing a generative adversarial network to generate spurious features.
Outcome: The proposed method is highly effective and adaptable on the ACE 2005 and TAC-KBP 2015 corpora.
Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models can be used to attack content filtering algorithms in social media platforms.
Approach: They propose to generate adversarial examples to test the robustness of social media content filtering algorithms.
Outcome: The proposed model outperforms existing models in the case of propaganda, false claims, rumours and hyperpartisan news.
On the Transferability of Adversarial Attacks against Neural Text Classifier (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that deep neural networks are vulnerable to adversarial examples . a small perturbation to an input alters the model prediction .
Approach: They propose a genetic algorithm to find models that can induce adversarial examples to fool models . they propose word replacement rules that can be used for model diagnostics from these examples .
Outcome: The proposed model can fool almost all existing models, while ignoring the data bias in the training set.
Detecting and Reducing Bias in a High Stakes Domain (D19-1)

Copied to clipboard

Challenge: Existing research shows that a deep learning model can predict aggression and loss in posts by focusing on stop words such as “a” or “on”.
Approach: They developed an approach to interpret a deep learning model that often bases its predictions on stop words such as "a" or "on" to tackle bias, they annotated the rationales and built models that drastically reduce bias.
Outcome: The proposed model can predict aggression and loss in posts by using stop words such as "a" or "on" the new annotations enable us to quantitatively measure how justified the model predictions are, and build models that drastically reduce bias.
KNOW How to Make Up Your Mind! Adversarially Detecting and Alleviating Inconsistencies in Natural Language Explanations (2023.acl-short)

Copied to clipboard

Challenge: eIA is an adversarial attack that generates inconsistent natural language explanations (NLEs) a model that generate In-NLE is undesirable, as it has a faulty decision-making process or is prone to inconsistencies.
Approach: They propose an off-the-shelf mitigation method to alleviate inconsistencies by grounding the model into external background knowledge.
Outcome: The proposed method reduces inconsistencies detected by previous models . it is based on external knowledge bases and a novel approach to mitigate inconsistent models based upon the proposed method .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations