Papers by Sujan Dutta
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive (2023.emnlp-main)
Copied to clipboard
Tharindu Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur KhudaBukhsh
| Challenge: | a paper examines how machine and human moderators disagree on offensive speech . offensive speech detection is a key component of content moderation . |
| Approach: | They propose a large-scale noise audit and a vicarious offense dataset to investigate disagreement on social web political discourse. |
| Outcome: | The proposed dataset reveals that moderation outcomes vary wildly across different machine moderators. |
What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs’ Self-consistency in Closed Domains Via Adversarial Nudge (2026.acl-long)
Copied to clipboard
| Challenge: | Claude exhibits strong resilience, while GPT and Grok demonstrate moderate resilience . open models fall short significantly, while proprietary models exhibit weak resilience compared to open models . |
| Approach: | They propose a framework for stress testing factual fidelity in large language models in the presence of adversarial nudges. |
| Outcome: | The proposed model is robust to adversarial nudges in two closed domains. |
Rater Cohesion and Quality from a Vicarious Perspective (2024.findings-emnlp)
Copied to clipboard
Deepak Pandita, Tharindu Cyril Weerasooriya, Sujan Dutta, Sarah Luger, Tharindu Ranasinghe, Ashiqur KhudaBukhsh, Marcos Zampieri, Christopher Homan
| Challenge: | Recent work in reinforcement learning with human feedback (RLHF) highlights the gains in model performance from aligning them to human values. |
| Approach: | They propose to use vicarious annotation to break down disagreement by asking raters how they think others would annotate the data. |
| Outcome: | The proposed method breaks down disagreements by asking raters how they think others would annotate the data. |