Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing systems that target hate speech with intent-conditioned counterspeech generate better results with longer contexts. |
| Approach: | They propose a framework that enables counterspeech generation by modeling the pragmatic implications underlying social biases in hateful statements. |
| Outcome: | The proposed framework outperforms existing benchmarks in intent-conditioned counterspeech generation. |
Similar Papers
Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech Generation (2023.acl-long)
Copied to clipboard
| Challenge: | a counterspeech with a certain intent may not be sufficient in every situation due to complex nature of hate speech . a novel framework for intent-conditioned counterseech generation is proposed to address the pervasive issue of hateful speech on the internet. |
| Approach: | They propose a framework for intent-conditioned counterspeech generation that leverages intent-specific representations and a fusion module to incorporate intent-related information into the model. |
| Outcome: | The proposed framework outperforms baselines by 10% across evaluation metrics. |
Outcome-Constrained Large Language Models for Countering Hate Speech (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven. |
| Approach: | They develop automatic counterspeech generation methods that incorporate two desired conversation outcomes into the text generation process: low conversation incivility and non-hateful hater reentry. |
| Outcome: | The proposed methods incorporate two desired conversation outcomes: low conversation incivility and non-hateful hater reentry. |
Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods to generate counterspeech based on intents are limited to single attributed . however, a holistic approach that considers multiple attributes simultaneously yields more nuanced and effective responses. |
| Approach: | They propose a framework that leverages hierarchical prefix learning with preference optimization to generate more constructive counterspeech. |
| Outcome: | The proposed framework improves intent conformity and emotion labels in 13,973 counterspeech instances. |
Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech (2021.findings-acl)
Copied to clipboard
| Challenge: | Off-the-shelf methods to generate hate speech are limited in that they generate repetitive and safe responses regardless of the hate speech. |
| Approach: | They propose a three-module pipeline approach to generate diverse and relevant counterspeech . they first generate various counterspeak candidates by a generative model, then filter ungrammatical ones using a BERT model . |
| Outcome: | The proposed pipeline generates diverse and relevant counterspeech responses on three datasets. |
F2RL: Factuality and Faithfulness Reinforcement Learning Framework for Claim-Guided Evidence-Supported Counterspeech Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for generating evidence-supported counterspeech lack clear guidance with a core claim for organizing evidence. |
| Approach: | They propose a Factuality and Faithfulness Reinforcement Learning framework for generating claim-guided and evidence-supported counterspeech (F2RL) they generate counter-claims based on hate speech and design a self-evaluation mechanism to select the most appropriate one. |
| Outcome: | The proposed framework achieves excellent performance on three benchmark datasets with strong factuality and faithfulness. |
Intent-Aware and Hate-Mitigating Counterspeech Generation via Dual-Discriminator Guided LLMs (2024.lrec-main)
Copied to clipboard
| Challenge: | Hate speech is an aggressive expression that incites hatred towards specific groups based on their group identity. |
| Approach: | They propose an LLMs-based framework for counterspeech generation that uses intent-aware discriminators to decode intents of LLM models. |
| Outcome: | The proposed framework matches intents with hate mitigation intents and performs well. |
Assessing the Human Likeness of AI-Generated Counterspeech (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies have focused on relevance, surface form, and other shallow linguistic characteristics. |
| Approach: | They propose to evaluate the human likeness of AI-generated counterspeech . they implement and evaluate several LLM-based generation strategies . |
| Outcome: | The proposed models show that human-written counterspeech can be distinguished by both simple classifiers and humans. |
Echoes of Discord: Forecasting Hater Reactions to Counterspeech (2025.findings-naacl)
Copied to clipboard
| Challenge: | Hate speech (HS) online causes increased prejudice and discrimination, fostering an environment of hostility and social division. |
| Approach: | They analyze the Reddit Echoes of Hate dataset to assess the impact of counterspeech from the hater's perspective and focus on whether the counterspeak leads the reentry to be hateful. |
| Outcome: | The proposed model outperforms the two-stage reaction predictor and the three-way classifier to predict haters' reactions to the reentry of the conversation and determines the type of resentment. |
Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Counterspeech, i.e. responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship risks of deletion-based content moderation. |
| Approach: | They draw from psychology and philosophy literature to craft six psychologically inspired strategies to challenge the underlying stereotypical implications of hateful language. |
| Outcome: | The strategies used in human- and machine-generated counterspeech datasets are convincing, whereas human-written counterspech uses less specific strategies compared to machine-produced counters. |
Counterspeech Generation using Small Language Models (2026.acl-srw)
Copied to clipboard
| Challenge: | Social media use is growing annually with about 68.5% of the global population active on these platforms as of July 2025. |
| Approach: | They evaluate SLMs ranging from 100 million to 3 billion parameters using simple prompting strategies as well as fine-tuning, combining automatic and robust human evaluations. |
| Outcome: | The proposed models generate relevant, coherent, and high-quality counterspeech, suggesting their suitability for efficient and responsible deployments. |