Papers by Sahajpreet Singh
Hate Personified: Investigating the role of LLMs in content moderation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Our work provides preliminary guidelines and highlights the nuances of applying Large Language models in culturally sensitive cases. |
| Approach: | They propose to use large language models to help with content moderation to assess how well the needs of diverse groups are reflected in annotated posts. |
| Outcome: | The proposed model is able to leverage community-based flagging efforts and exposure to adversaries. |
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing systems that target hate speech with intent-conditioned counterspeech generate better results with longer contexts. |
| Approach: | They propose a framework that enables counterspeech generation by modeling the pragmatic implications underlying social biases in hateful statements. |
| Outcome: | The proposed framework outperforms existing benchmarks in intent-conditioned counterspeech generation. |