Validating Automatic Evaluation of Controllable Counterspeech Generation: Rankings Matter More Than Scores (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing methods for evaluating attributes of counterspeech are limited and the validity of such evaluations is questionable when the classifiers themselves have only modest performance. |
| Approach: | They examine the automatic evaluation of counterspeech attributes using a multi-attribute counterseech dataset containing 2,728 samples. |
| Outcome: | The proposed model can be trusted by classifier validation, and it can rank models with confidence. |
Similar Papers
Assessing the Human Likeness of AI-Generated Counterspeech (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies have focused on relevance, surface form, and other shallow linguistic characteristics. |
| Approach: | They propose to evaluate the human likeness of AI-generated counterspeech . they implement and evaluate several LLM-based generation strategies . |
| Outcome: | The proposed models show that human-written counterspeech can be distinguished by both simple classifiers and humans. |
CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs (2025.naacl-long)
Copied to clipboard
| Challenge: | Current evaluation methods do not capture complex attributes of counterspeech quality, such as contextual relevance, aggressiveness, or argumentative coherence. |
| Approach: | They propose to use a dataset and framework to evaluate counterspeech quality across four dimensions: contextual relevance, aggressiveness, argument-coherence, and suitability. |
| Outcome: | The proposed method outperforms ROUGE, METEOR, and BertScore in correlating with human judgement, indicating a significant improvement in automated counterspeech evaluation. |
Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods to generate counterspeech based on intents are limited to single attributed . however, a holistic approach that considers multiple attributes simultaneously yields more nuanced and effective responses. |
| Approach: | They propose a framework that leverages hierarchical prefix learning with preference optimization to generate more constructive counterspeech. |
| Outcome: | The proposed framework improves intent conformity and emotion labels in 13,973 counterspeech instances. |
NLP for Counterspeech against Hate: A Survey and How-To Guide (2024.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have focused on the challenges of analysing, collecting, classifying, and automatically generating counterspeech, to reduce the huge burden of manually producing it. |
| Approach: | They propose a guide for doing research on counterspeech, with detailed examples and best practices that can be learnt from the NLP community. |
| Outcome: | The proposed strategies can reduce online and offline violence while preserving the freedom of speech of the users. |
Counterspeech Generation using Small Language Models (2026.acl-srw)
Copied to clipboard
| Challenge: | Social media use is growing annually with about 68.5% of the global population active on these platforms as of July 2025. |
| Approach: | They evaluate SLMs ranging from 100 million to 3 billion parameters using simple prompting strategies as well as fine-tuning, combining automatic and robust human evaluations. |
| Outcome: | The proposed models generate relevant, coherent, and high-quality counterspeech, suggesting their suitability for efficient and responsible deployments. |
On the Effectiveness of Adversarial Robustness for Abuse Mitigation with Counterspeech (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent work on automated counterspeech systems focused on synthetic data but rarely looked into how the public deals with abuse. |
| Approach: | They propose to curate a new dataset of abuse and replies from footballers for study of public figure abuse and use it to examine how models can handle adversarial attacks. |
| Outcome: | The proposed model is robust against adversarial attacks across domains and can handle abuse in the real world. |
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing systems that target hate speech with intent-conditioned counterspeech generate better results with longer contexts. |
| Approach: | They propose a framework that enables counterspeech generation by modeling the pragmatic implications underlying social biases in hateful statements. |
| Outcome: | The proposed framework outperforms existing benchmarks in intent-conditioned counterspeech generation. |
Outcome-Constrained Large Language Models for Countering Hate Speech (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven. |
| Approach: | They develop automatic counterspeech generation methods that incorporate two desired conversation outcomes into the text generation process: low conversation incivility and non-hateful hater reentry. |
| Outcome: | The proposed methods incorporate two desired conversation outcomes: low conversation incivility and non-hateful hater reentry. |
Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Counterspeech, i.e. responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship risks of deletion-based content moderation. |
| Approach: | They draw from psychology and philosophy literature to craft six psychologically inspired strategies to challenge the underlying stereotypical implications of hateful language. |
| Outcome: | The strategies used in human- and machine-generated counterspeech datasets are convincing, whereas human-written counterspech uses less specific strategies compared to machine-produced counters. |
Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech Generation (2023.acl-long)
Copied to clipboard
| Challenge: | a counterspeech with a certain intent may not be sufficient in every situation due to complex nature of hate speech . a novel framework for intent-conditioned counterseech generation is proposed to address the pervasive issue of hateful speech on the internet. |
| Approach: | They propose a framework for intent-conditioned counterspeech generation that leverages intent-specific representations and a fusion module to incorporate intent-related information into the model. |
| Outcome: | The proposed framework outperforms baselines by 10% across evaluation metrics. |