Papers by Sujan Dutta

3 papers
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive (2023.emnlp-main)

Copied to clipboard

Challenge: a paper examines how machine and human moderators disagree on offensive speech . offensive speech detection is a key component of content moderation .
Approach: They propose a large-scale noise audit and a vicarious offense dataset to investigate disagreement on social web political discourse.
Outcome: The proposed dataset reveals that moderation outcomes vary wildly across different machine moderators.
What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs’ Self-consistency in Closed Domains Via Adversarial Nudge (2026.acl-long)

Copied to clipboard

Challenge: Claude exhibits strong resilience, while GPT and Grok demonstrate moderate resilience . open models fall short significantly, while proprietary models exhibit weak resilience compared to open models .
Approach: They propose a framework for stress testing factual fidelity in large language models in the presence of adversarial nudges.
Outcome: The proposed model is robust to adversarial nudges in two closed domains.
Rater Cohesion and Quality from a Vicarious Perspective (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent work in reinforcement learning with human feedback (RLHF) highlights the gains in model performance from aligning them to human values.
Approach: They propose to use vicarious annotation to break down disagreement by asking raters how they think others would annotate the data.
Outcome: The proposed method breaks down disagreements by asking raters how they think others would annotate the data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations