Challenge: Existing methods for evaluating attributes of counterspeech are limited and the validity of such evaluations is questionable when the classifiers themselves have only modest performance.
Approach: They examine the automatic evaluation of counterspeech attributes using a multi-attribute counterseech dataset containing 2,728 samples.
Outcome: The proposed model can be trusted by classifier validation, and it can rank models with confidence.

Similar Papers

Assessing the Human Likeness of AI-Generated Counterspeech (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have focused on relevance, surface form, and other shallow linguistic characteristics.
Approach: They propose to evaluate the human likeness of AI-generated counterspeech . they implement and evaluate several LLM-based generation strategies .
Outcome: The proposed models show that human-written counterspeech can be distinguished by both simple classifiers and humans.
CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Current evaluation methods do not capture complex attributes of counterspeech quality, such as contextual relevance, aggressiveness, or argumentative coherence.
Approach: They propose to use a dataset and framework to evaluate counterspeech quality across four dimensions: contextual relevance, aggressiveness, argument-coherence, and suitability.
Outcome: The proposed method outperforms ROUGE, METEOR, and BertScore in correlating with human judgement, indicating a significant improvement in automated counterspeech evaluation.
Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to generate counterspeech based on intents are limited to single attributed . however, a holistic approach that considers multiple attributes simultaneously yields more nuanced and effective responses.
Approach: They propose a framework that leverages hierarchical prefix learning with preference optimization to generate more constructive counterspeech.
Outcome: The proposed framework improves intent conformity and emotion labels in 13,973 counterspeech instances.
NLP for Counterspeech against Hate: A Survey and How-To Guide (2024.findings-naacl)

Copied to clipboard

Challenge: Recent studies have focused on the challenges of analysing, collecting, classifying, and automatically generating counterspeech, to reduce the huge burden of manually producing it.
Approach: They propose a guide for doing research on counterspeech, with detailed examples and best practices that can be learnt from the NLP community.
Outcome: The proposed strategies can reduce online and offline violence while preserving the freedom of speech of the users.
Counterspeech Generation using Small Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Social media use is growing annually with about 68.5% of the global population active on these platforms as of July 2025.
Approach: They evaluate SLMs ranging from 100 million to 3 billion parameters using simple prompting strategies as well as fine-tuning, combining automatic and robust human evaluations.
Outcome: The proposed models generate relevant, coherent, and high-quality counterspeech, suggesting their suitability for efficient and responsible deployments.
On the Effectiveness of Adversarial Robustness for Abuse Mitigation with Counterspeech (2024.naacl-long)

Copied to clipboard

Challenge: Recent work on automated counterspeech systems focused on synthetic data but rarely looked into how the public deals with abuse.
Approach: They propose to curate a new dataset of abuse and replies from footballers for study of public figure abuse and use it to examine how models can handle adversarial attacks.
Outcome: The proposed model is robust against adversarial attacks across domains and can handle abuse in the real world.
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF (2024.naacl-long)

Copied to clipboard

Challenge: Existing systems that target hate speech with intent-conditioned counterspeech generate better results with longer contexts.
Approach: They propose a framework that enables counterspeech generation by modeling the pragmatic implications underlying social biases in hateful statements.
Outcome: The proposed framework outperforms existing benchmarks in intent-conditioned counterspeech generation.
Outcome-Constrained Large Language Models for Countering Hate Speech (2024.emnlp-main)

Copied to clipboard

Challenge: Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven.
Approach: They develop automatic counterspeech generation methods that incorporate two desired conversation outcomes into the text generation process: low conversation incivility and non-hateful hater reentry.
Outcome: The proposed methods incorporate two desired conversation outcomes: low conversation incivility and non-hateful hater reentry.
Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language (2023.findings-emnlp)

Copied to clipboard

Challenge: Counterspeech, i.e. responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship risks of deletion-based content moderation.
Approach: They draw from psychology and philosophy literature to craft six psychologically inspired strategies to challenge the underlying stereotypical implications of hateful language.
Outcome: The strategies used in human- and machine-generated counterspeech datasets are convincing, whereas human-written counterspech uses less specific strategies compared to machine-produced counters.
Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech Generation (2023.acl-long)

Copied to clipboard

Challenge: a counterspeech with a certain intent may not be sufficient in every situation due to complex nature of hate speech . a novel framework for intent-conditioned counterseech generation is proposed to address the pervasive issue of hateful speech on the internet.
Approach: They propose a framework for intent-conditioned counterspeech generation that leverages intent-specific representations and a fusion module to incorporate intent-related information into the model.
Outcome: The proposed framework outperforms baselines by 10% across evaluation metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations