Papers by Kshitiz Tiwari

1 papers
Robust Hate Speech Detection via Mitigating Spurious Correlations (2022.aacl-short)

Copied to clipboard

Challenge: a novel hate speech detection model can be used to detect word- and character-level adversarial attacks . existing adversarials assume that attackers replace the target words with other names to evade detection .
Approach: They propose a robust hate speech detection model that can defend against adversarial attacks . they describe the process of hate speech recognition by a causal graph and a regularized entropy loss function to quantify spurious correlation .
Outcome: The proposed model can defend against word- and character-level adversarial attacks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations