Papers by Tharindu Weerasooriya
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive (2023.emnlp-main)
Copied to clipboard
Tharindu Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur KhudaBukhsh
| Challenge: | a paper examines how machine and human moderators disagree on offensive speech . offensive speech detection is a key component of content moderation . |
| Approach: | They propose a large-scale noise audit and a vicarious offense dataset to investigate disagreement on social web political discourse. |
| Outcome: | The proposed dataset reveals that moderation outcomes vary wildly across different machine moderators. |
Rater Cohesion and Quality from a Vicarious Perspective (2024.findings-emnlp)
Copied to clipboard
Deepak Pandita, Tharindu Cyril Weerasooriya, Sujan Dutta, Sarah Luger, Tharindu Ranasinghe, Ashiqur KhudaBukhsh, Marcos Zampieri, Christopher Homan
| Challenge: | Recent work in reinforcement learning with human feedback (RLHF) highlights the gains in model performance from aligning them to human values. |
| Approach: | They propose to use vicarious annotation to break down disagreement by asking raters how they think others would annotate the data. |
| Outcome: | The proposed method breaks down disagreements by asking raters how they think others would annotate the data. |