Challenge: Existing methods to generalize hate speech detection models have been limited by the labeling criteria between datasets.
Approach: They propose a framework that uses the concept of multi-agent for hate speech detection that uses a set of labeling criteria to create multiple agents based on the induced labeling of given datasets.
Outcome: The proposed framework achieves superior cross-evaluation performance compared to methods that focus on specific labeling criteria or majority voting methods.

Similar Papers

Debate-Feedback: A Multi-Agent Framework for Efficient Legal Judgment Prediction (2025.naacl-short)

Copied to clipboard

Challenge: Comparative experiments show that our model outperforms several general-purpose and domain-specific legal models.
Approach: They propose a legal judgment prediction model that integrates LLMs with argumentative reasoning techniques to simulate the debate phase of real courtroom trials.
Outcome: The proposed model outperforms several general-purpose and domain-specific legal models and offers a dynamic reasoning process.
More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection (2026.findings-acl)

Copied to clipboard

Challenge: Existing systems struggle with multimodal content where the emergent meaning transcends the aggregation of individual modalities.
Approach: They propose a framework to characterize semantic intent shifts where modalities interact to construct implicit hate from benign cues or neutralize toxicity through semantic inversion.
Outcome: The proposed framework outperforms state-of-the-art benchmarks on H-VLI and on established benchmarks.
What Did You Learn To Hate? A Topic-Oriented Analysis of Generalization in Hate Speech Detection (2023.eacl-main)

Copied to clipboard

Challenge: Hate speech detection datasets often use different annotation guidelines, resulting in inconsistencies . authors propose a topic-oriented approach to study generalization across popular hate speech datasets .
Approach: They propose a topic-oriented approach to study generalization across popular hate speech datasets . they compare Transformer-based models in capturing topic-generic and topic-specific knowledge .
Outcome: The proposed approach improves the reliability of hate speech detection on social media platforms.
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection (2026.acl-long)

Copied to clipboard

Challenge: a new framework for hate speech detection addresses implicit hate speech by tailoring the detection process to dataset-specific attributes.
Approach: They propose a framework to account for the dataset-specific characteristics of hate speech datasets.
Outcome: The proposed framework improves detection accuracy and provides interpretable insights into the distinctive features of each dataset.
A Benchmark Dataset for Learning to Intervene in Online Hate Speech (D19-1)

Copied to clipboard

Challenge: Existing methods to detect online hate speech ignore conversational context . generative hate speech intervention is a novel approach to counter online hate .
Approach: They propose a task where generative hate speech intervention generates responses to intervene during online conversations that contain hate speech.
Outcome: The proposed method can detect and block hate speech and discourage it . it can also generate responses written by Mechanical Turk workers .
Fight Fire with Fire: Fine-tuning Hate Detectors using Large Samples of Generated Hate Speech (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for hate speech detection are limited in size and lack of labeled datasets.
Approach: They employ pretrained language models to generate large amounts of hate speech sequences from available labeled examples.
Outcome: The proposed model improves generalization significantly and consistently within and across data distributions.
Improving Hate Speech Detection with Deep Learning Ensembles (L18-1)

Copied to clipboard

Challenge: censorship is a potential risk when addressing these issues with automated text classification methods.
Approach: They propose to use a neural network-based ensemble method to better classify hate speech using a publicly available embedding model and a popular sentiment dataset.
Outcome: The proposed method improves by 5 points on a hate speech corpus from Twitter and a popular sentiment dataset.
SharedCon: Implicit Hate Speech Detection using Shared Semantics (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies suggest that classifying hateful posts in a binary manner may not address nuanced task of detecting implicit hate speech.
Approach: They propose a contrastive learning approach that leverages shared semantics among data to detect implicit hate speech.
Outcome: The proposed approach is based on a clustering-based contrastive learning approach with human-written implications or machine-generated augmented data.
A Federated Approach for Hate Speech Detection (2023.eacl-main)

Copied to clipboard

Challenge: Despite the scale of social media content, privacy preservation in hate speech detection has remained understudied.
Approach: They propose to use federated machine learning to address privacy concerns in hate speech detection by obtaining a 6.81% improvement in F1-score.
Outcome: The proposed method improves the F1-score of hate speech detection by 6.81% while maintaining public data privacy.
Mitigating Biases in Hate Speech Detection from A Causal Perspective (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to detect hate speech are prone to spurious correlations between training data and labels, which could lead to biased treatment of vulnerable and minority groups.
Approach: They propose to use grammar induction to find grammar patterns for hate speech and analyze this phenomenon from a causal perspective.
Outcome: The proposed methods can detect hate speech from a causal perspective and are robust to different datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations