Papers by Berk Atıl
♪ Something Just Like TRuST ♪ *: Toxicity Recognition of Span and Target (2026.findings-acl)
Copied to clipboard
| Challenge: | Toxic language is pervasive online, and because LLMs are trained on web data, it generates such content. |
| Approach: | They propose a large-scale dataset that synthesizes toxicity definitions and an annotation scheme . they use a rigorous human annotation process to evaluate the diversity of the annotations . |
| Outcome: | The proposed model outperforms existing models on three tasks and is not reliable. |
Self-Explaining Hate Speech Detection with Moral Rationales (2026.findings-acl)
Copied to clipboard
Francielle Vargas, Jackson Trager, Diego Alves, Matteo Guida, Surendrabikram Thapa, Berk Atıl, Daryna Dementieva, Andrew J Smart, Ameeta Agrawal
| Challenge: | Existing models for hate speech detection are opaque and rely on surface-level cues. Existing approaches often encode biases originating from training data and annotation processes. |
| Approach: | They propose a framework that integrates moral rationale supervision into training . they propose SMRA for self-explaining hate speech detection . |
| Outcome: | The proposed framework improves performance across binary hate speech detection and multi-label moral sentiment classification. |