Papers by Kshitiz Tiwari
Robust Hate Speech Detection via Mitigating Spurious Correlations (2022.aacl-short)
Copied to clipboard
| Challenge: | a novel hate speech detection model can be used to detect word- and character-level adversarial attacks . existing adversarials assume that attackers replace the target words with other names to evade detection . |
| Approach: | They propose a robust hate speech detection model that can defend against adversarial attacks . they describe the process of hate speech recognition by a causal graph and a regularized entropy loss function to quantify spurious correlation . |
| Outcome: | The proposed model can defend against word- and character-level adversarial attacks. |