| Challenge: | Despite the scale of social media content, privacy preservation in hate speech detection has remained understudied. |
| Approach: | They propose to use federated machine learning to address privacy concerns in hate speech detection by obtaining a 6.81% improvement in F1-score. |
| Outcome: | The proposed method improves the F1-score of hate speech detection by 6.81% while maintaining public data privacy. |
Similar Papers
Privacy-Preserving Federated Learning for Hate Speech Detection (2025.naacl-srw)
Copied to clipboard
| Challenge: | a federated learning system with differential privacy is tailored to low-resource languages . data with fewer than 20 sentences per client struggled due to excessive noise . |
| Approach: | They propose a federated learning system with differential privacy for hate speech detection . they fine-tuned pre-trained language models to find it to be the most effective . |
| Outcome: | The proposed learning system outperforms other models in low-resource languages . balanced datasets and augmenting hateful data with non-hateful examples proved critical . |
Improving Hate Speech Detection with Deep Learning Ensembles (L18-1)
Copied to clipboard
| Challenge: | censorship is a potential risk when addressing these issues with automated text classification methods. |
| Approach: | They propose to use a neural network-based ensemble method to better classify hate speech using a publicly available embedding model and a popular sentiment dataset. |
| Outcome: | The proposed method improves by 5 points on a hate speech corpus from Twitter and a popular sentiment dataset. |
Generalizable Multilingual Hate Speech Detection on Low Resource Indian Languages using Fair Selection in Federated Learning (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for detecting hate speech in Indian languages with linguistic diversity and cultural nuances are undesirable and involving potential risk to their privacy. |
| Approach: | They propose a federated approach that utilizes continuous adaptation and fine-tuning to aid generalization using subsets of multilingual data. |
| Outcome: | The proposed approach outperforms the state-of-the-art models on 13 Indic datasets across five different pre-trained models. |
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection (2026.acl-long)
Copied to clipboard
| Challenge: | a new framework for hate speech detection addresses implicit hate speech by tailoring the detection process to dataset-specific attributes. |
| Approach: | They propose a framework to account for the dataset-specific characteristics of hate speech datasets. |
| Outcome: | The proposed framework improves detection accuracy and provides interpretable insights into the distinctive features of each dataset. |
Hate Speech Detection Based on Sentiment Knowledge Sharing (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods for hate speech detection are stereotyped and biased . et al., a paper examining the effectiveness of multitask learning in hate speech recognition tasks . |
| Approach: | They propose a hate speech detection framework based on sentiment knowledge sharing . they extract affective features of the target sentence and use sentiment features from external resources . |
| Outcome: | The proposed model can detect hate speech over two public datasets. |
Mitigating Biases in Hate Speech Detection from A Causal Perspective (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to detect hate speech are prone to spurious correlations between training data and labels, which could lead to biased treatment of vulnerable and minority groups. |
| Approach: | They propose to use grammar induction to find grammar patterns for hate speech and analyze this phenomenon from a causal perspective. |
| Outcome: | The proposed methods can detect hate speech from a causal perspective and are robust to different datasets. |
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter (2025.acl-long)
Copied to clipboard
Manuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale, Samuel Fraiberger, Victor Orozco-Olvera, Paul Röttger
| Challenge: | Prior work on automated hate speech detection models has been limited due to systematic biases in evaluation datasets and poor performance across geographies. |
| Approach: | They propose to construct a global hate speech dataset representative of social media settings from tweets posted on September 21, 2022. |
| Outcome: | The proposed dataset covers eight languages and four English-speaking countries and covers eight countries where English is the main language on Twitter. |
FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data (2026.acl-long)
Copied to clipboard
| Challenge: | Social media text data is often used to train machine learning models to identify users exhibiting high-risk mental health behaviors. |
| Approach: | They apply federatedlearning and Differentially Private FL to two widely-studied mental health prediction tasks using social media text data. |
| Outcome: | The proposed methods achieve comparable performance to centralized training on depression identification, but have a large performance-privacy trade-off even with low levels of noise. |
Data-Efficient Methods For Improving Hate Speech Detection (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods for hate speech detection are data-hungry and require large datasets. |
| Approach: | They propose an input-level data augmentation technique EasyMix to improve hate speech detection in english and multilingual datasets. |
| Outcome: | The proposed method improves the performance across english and multilingual datasets by 1% and 2-8%. |
Darkness can not drive out darkness: Investigating Bias in Hate SpeechDetection Models (2022.acl-srw)
Copied to clipboard
| Challenge: | a recent study shows that machine learning models are biased and they might make the wrong decisions for the wrong reasons. |
| Approach: | They investigate the impact of social bias on the performance of hate speech detection models . they also investigate the causal effect of intersectional bias on models' unfairness . |
| Outcome: | The proposed model is biased and makes the wrong decisions for the wrong reasons. |