Papers by Enes Altinisik
Impact of Adversarial Training on Robustness and Generalizability of Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Adversarial training is widely acknowledged as the most effective defense against adversarial attacks, but achieving both robustness and generalization requires a trade-off. |
| Approach: | They propose to compare pre-training data augmentation and training time input perturbations with embedding space perturbations to find out whether they improve generalization. |
| Outcome: | The proposed methods improve generalization and robustness of the trained models. |
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Current content moderation filters focus on general safety and ignore cultural context . authors: FanarGuard improves accuracy and provides a practical step toward context-sensitive safeguards. |
| Approach: | They propose a bilingual moderation filter that evaluates both safety and cultural alignment in Arabic and English. |
| Outcome: | The proposed moderation filter performs better with human annotations than state-of-the-art filters on safety benchmarks. |