Papers by Urja Khurana
DefVerify: Do Hate Speech Models Reflect Their Dataset’s Definition? (2025.coling-main)
Copied to clipboard
| Challenge: | DefVerify is a 3-step procedure that encodes a user-specified definition of hate speech, quantifies to what extent the model reflects the intended definition, and identifies the point of failure in the workflow. |
| Approach: | They propose a 3-step procedure that encodes a user-specified definition of hate speech and quantifies to what extent the model reflects intended definition. |
| Outcome: | The proposed procedure detects gaps between definition and model behavior when applied to six popular hate speech benchmark datasets. |