Challenge: Automated responses lack argumentative richness which characterises expert-produced counterspeech.
Approach: They propose to automate counterspeech generation by investigating tension between helpfulness and harmlessness of LLMs and to assess whether presence of safety guardrails hinders quality of generations.
Outcome: The proposed approach produces more cogent responses that lack argumentative richness which characterises expert-produced counterspeech.

Similar Papers

Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech (2021.findings-acl)

Copied to clipboard

Challenge: Off-the-shelf methods to generate hate speech are limited in that they generate repetitive and safe responses regardless of the hate speech.
Approach: They propose a three-module pipeline approach to generate diverse and relevant counterspeech . they first generate various counterspeak candidates by a generative model, then filter ungrammatical ones using a BERT model .
Outcome: The proposed pipeline generates diverse and relevant counterspeech responses on three datasets.
NLP for Counterspeech against Hate: A Survey and How-To Guide (2024.findings-naacl)

Copied to clipboard

Challenge: Recent studies have focused on the challenges of analysing, collecting, classifying, and automatically generating counterspeech, to reduce the huge burden of manually producing it.
Approach: They propose a guide for doing research on counterspeech, with detailed examples and best practices that can be learnt from the NLP community.
Outcome: The proposed strategies can reduce online and offline violence while preserving the freedom of speech of the users.
NLP for Counterspeech against Hate and Misinformation (CSHAM) (2025.acl-tutorials)

Copied to clipboard

Challenge: tutorial aims to show how counterspeech is used to tackle abuse and misinformation by individuals, activists and organisations.
Approach: tutorial aims to show how counterspeech is currently used to tackle abuse and misinformation . will also show how Natural Language Processing (NLP) and Generation (NLG) can be applied to automate its production.
Outcome: The tutorial will bring diverse multidisciplinary perspectives to safety research . case studies from industry and public policy will be included .
Guardrails and Security for LLMs: Safe, Secure and Controllable Steering of LLM Applications (2025.acl-tutorials)

Copied to clipboard

Challenge: Pretrained generative models provide novel ways for users to interact with computers.
Approach: This tutorial provides an overview of key guardrail mechanisms developed for LLMs along with evaluation methodologies and a detailed security assessment protocol.
Outcome: This tutorial provides an overview of key guardrail mechanisms developed for LLMs, along with evaluation methodologies and a detailed security assessment protocol.
Outcome-Constrained Large Language Models for Countering Hate Speech (2024.emnlp-main)

Copied to clipboard

Challenge: Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven.
Approach: They develop automatic counterspeech generation methods that incorporate two desired conversation outcomes into the text generation process: low conversation incivility and non-hateful hater reentry.
Outcome: The proposed methods incorporate two desired conversation outcomes: low conversation incivility and non-hateful hater reentry.
LLM generated responses to mitigate the impact of hate speech (2024.findings-emnlp)

Copied to clipboard

Challenge: a study aims to determine the effectiveness of large language models to counteract hate speech . it is the first real-life A/B test evaluating the effectiveness .
Approach: They conduct the first real-life A/B test assessing the effectiveness of LLM-generated counter-speech.
Outcome: The proposed system reduces user engagement by over 20%, the study shows . the proposed metric is based on a simple metric and is scalable to other platforms .
Assessing the Human Likeness of AI-Generated Counterspeech (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have focused on relevance, surface form, and other shallow linguistic characteristics.
Approach: They propose to evaluate the human likeness of AI-generated counterspeech . they implement and evaluate several LLM-based generation strategies .
Outcome: The proposed models show that human-written counterspeech can be distinguished by both simple classifiers and humans.
Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language (2023.findings-emnlp)

Copied to clipboard

Challenge: Counterspeech, i.e. responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship risks of deletion-based content moderation.
Approach: They draw from psychology and philosophy literature to craft six psychologically inspired strategies to challenge the underlying stereotypical implications of hateful language.
Outcome: The strategies used in human- and machine-generated counterspeech datasets are convincing, whereas human-written counterspech uses less specific strategies compared to machine-produced counters.
Counterspeech Generation using Small Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Social media use is growing annually with about 68.5% of the global population active on these platforms as of July 2025.
Approach: They evaluate SLMs ranging from 100 million to 3 billion parameters using simple prompting strategies as well as fine-tuning, combining automatic and robust human evaluations.
Outcome: The proposed models generate relevant, coherent, and high-quality counterspeech, suggesting their suitability for efficient and responsible deployments.
Characterizing Selective Refusal Bias in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: a recent study shows that safety guardrails in large language models can inadvertently introduce or reflect new biases as they may refuse to generate harmful content targeting some demographic groups and not others.
Approach: They examine the selective refusal bias in large language models by examining demographics and responses.
Outcome: The proposed model fails to defend against an indirect attack on previously refused groups in 89% of the trials.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations