A Federated Approach for Hate Speech Detection (2023.eacl-main)

Copied to clipboard

Challenge: Despite the scale of social media content, privacy preservation in hate speech detection has remained understudied.
Approach: They propose to use federated machine learning to address privacy concerns in hate speech detection by obtaining a 6.81% improvement in F1-score.
Outcome: The proposed method improves the F1-score of hate speech detection by 6.81% while maintaining public data privacy.

Similar Papers

Privacy-Preserving Federated Learning for Hate Speech Detection (2025.naacl-srw)

Copied to clipboard

Challenge: a federated learning system with differential privacy is tailored to low-resource languages . data with fewer than 20 sentences per client struggled due to excessive noise .
Approach: They propose a federated learning system with differential privacy for hate speech detection . they fine-tuned pre-trained language models to find it to be the most effective .
Outcome: The proposed learning system outperforms other models in low-resource languages . balanced datasets and augmenting hateful data with non-hateful examples proved critical .
Improving Hate Speech Detection with Deep Learning Ensembles (L18-1)

Copied to clipboard

Challenge: censorship is a potential risk when addressing these issues with automated text classification methods.
Approach: They propose to use a neural network-based ensemble method to better classify hate speech using a publicly available embedding model and a popular sentiment dataset.
Outcome: The proposed method improves by 5 points on a hate speech corpus from Twitter and a popular sentiment dataset.
Generalizable Multilingual Hate Speech Detection on Low Resource Indian Languages using Fair Selection in Federated Learning (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for detecting hate speech in Indian languages with linguistic diversity and cultural nuances are undesirable and involving potential risk to their privacy.
Approach: They propose a federated approach that utilizes continuous adaptation and fine-tuning to aid generalization using subsets of multilingual data.
Outcome: The proposed approach outperforms the state-of-the-art models on 13 Indic datasets across five different pre-trained models.
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection (2026.acl-long)

Copied to clipboard

Challenge: a new framework for hate speech detection addresses implicit hate speech by tailoring the detection process to dataset-specific attributes.
Approach: They propose a framework to account for the dataset-specific characteristics of hate speech datasets.
Outcome: The proposed framework improves detection accuracy and provides interpretable insights into the distinctive features of each dataset.
Hate Speech Detection Based on Sentiment Knowledge Sharing (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for hate speech detection are stereotyped and biased . et al., a paper examining the effectiveness of multitask learning in hate speech recognition tasks .
Approach: They propose a hate speech detection framework based on sentiment knowledge sharing . they extract affective features of the target sentence and use sentiment features from external resources .
Outcome: The proposed model can detect hate speech over two public datasets.
Mitigating Biases in Hate Speech Detection from A Causal Perspective (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to detect hate speech are prone to spurious correlations between training data and labels, which could lead to biased treatment of vulnerable and minority groups.
Approach: They propose to use grammar induction to find grammar patterns for hate speech and analyze this phenomenon from a causal perspective.
Outcome: The proposed methods can detect hate speech from a causal perspective and are robust to different datasets.
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter (2025.acl-long)

Copied to clipboard

Challenge: Prior work on automated hate speech detection models has been limited due to systematic biases in evaluation datasets and poor performance across geographies.
Approach: They propose to construct a global hate speech dataset representative of social media settings from tweets posted on September 21, 2022.
Outcome: The proposed dataset covers eight languages and four English-speaking countries and covers eight countries where English is the main language on Twitter.
FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data (2026.acl-long)

Copied to clipboard

Challenge: Social media text data is often used to train machine learning models to identify users exhibiting high-risk mental health behaviors.
Approach: They apply federatedlearning and Differentially Private FL to two widely-studied mental health prediction tasks using social media text data.
Outcome: The proposed methods achieve comparable performance to centralized training on depression identification, but have a large performance-privacy trade-off even with low levels of noise.
Data-Efficient Methods For Improving Hate Speech Detection (2023.findings-eacl)

Copied to clipboard

Challenge: Existing methods for hate speech detection are data-hungry and require large datasets.
Approach: They propose an input-level data augmentation technique EasyMix to improve hate speech detection in english and multilingual datasets.
Outcome: The proposed method improves the performance across english and multilingual datasets by 1% and 2-8%.
Darkness can not drive out darkness: Investigating Bias in Hate SpeechDetection Models (2022.acl-srw)

Copied to clipboard

Challenge: a recent study shows that machine learning models are biased and they might make the wrong decisions for the wrong reasons.
Approach: They investigate the impact of social bias on the performance of hate speech detection models . they also investigate the causal effect of intersectional bias on models' unfairness .
Outcome: The proposed model is biased and makes the wrong decisions for the wrong reasons.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations