Challenge: Existing approaches to detect hate speech are expensive and time-consuming . a new approach allows for flexible learning of neighborhood information .
Approach: They propose a method that allows flexible modeling of neighbors retrieved from a resource-rich corpus to learn the amount of transfer.
Outcome: The proposed training strategy improves on low-resource hate speech corpora over baselines.

Similar Papers

Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for detecting hate speech data are expensive and time-consuming . labeled data is expensive and difficult to collect, especially for low-resource languages .
Approach: They propose a method that leverages nearest-neighbor retrieval to augment minimal labeled data in target language.
Outcome: The proposed method outperforms existing models on eight languages and is highly data-efficient.
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Hate speech datasets focus on English-language content, hindering effective models . annotating hateful content is expensive, time-consuming and potentially harmful to annotators.
Approach: They propose to use ISO 639-1 codes to fine-tune models on one source language and apply them to another language.
Outcome: The proposed approach performs well on some tasks, but fails on many others.
A Neighborhood Framework for Resource-Lean Content Flagging (2022.tacl-1)

Copied to clipboard

Challenge: Existing approaches to cross-lingual content flagging with limited target language data are lacking in many languages.
Approach: They propose a framework for cross-lingual content flagging with limited target- language data based on a nearest-neighbor architecture and a transformer representation in all its components.
Outcome: The proposed framework outperforms previous work in terms of predictive performance on eight languages from two different datasets.
Data-Efficient Methods For Improving Hate Speech Detection (2023.findings-eacl)

Copied to clipboard

Challenge: Existing methods for hate speech detection are data-hungry and require large datasets.
Approach: They propose an input-level data augmentation technique EasyMix to improve hate speech detection in english and multilingual datasets.
Outcome: The proposed method improves the performance across english and multilingual datasets by 1% and 2-8%.
SharedCon: Implicit Hate Speech Detection using Shared Semantics (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies suggest that classifying hateful posts in a binary manner may not address nuanced task of detecting implicit hate speech.
Approach: They propose a contrastive learning approach that leverages shared semantics among data to detect implicit hate speech.
Outcome: The proposed approach is based on a clustering-based contrastive learning approach with human-written implications or machine-generated augmented data.
Bridging Modalities: Enhancing Cross-Modality Hate Speech Detection with Few-Shot In-Context Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has developed models targeting specific modalities but lacks transferability between formats.
Approach: They conduct extensive experiments using few-shot in-context learning with large language models to explore the transferability of hate speech detection between modalities.
Outcome: The proposed model outperforms vision-language demonstrations in few-shot learning settings.
Privacy-Preserving Federated Learning for Hate Speech Detection (2025.naacl-srw)

Copied to clipboard

Challenge: a federated learning system with differential privacy is tailored to low-resource languages . data with fewer than 20 sentences per client struggled due to excessive noise .
Approach: They propose a federated learning system with differential privacy for hate speech detection . they fine-tuned pre-trained language models to find it to be the most effective .
Outcome: The proposed learning system outperforms other models in low-resource languages . balanced datasets and augmenting hateful data with non-hateful examples proved critical .
Improving Hate Speech Detection by Fusing Textual and User Interaction Representations in Online Communities (2026.acl-industry)

Copied to clipboard

Challenge: Existing studies on toxic content in online communities are limited by the scarcity of data that align textual content with comprehensive social interactions.
Approach: They propose a user-aware hate speech detection framework that effectively fuses textual semantics with social interaction representations to provide pragmatic context for disambiguation.
Outcome: The proposed framework outperforms strong text-only baselines by over 3.6%, validating the critical role of social context in enhancing detection accuracy.
Large-Scale Hate Speech Detection with Cross-Domain Transfer (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for hate speech detection are limited due to the labor cost.
Approach: They construct large-scale tweet datasets for hate speech detection in English and a low-resource language, Turkish, consisting of human-labeled 100k tweets per each.
Outcome: The proposed datasets outperform conventional bag-of-words and neural models by at least 5% in English and 10% in Turkish for large-scale hate speech detection.
Dynamically Refined Regularization for Improving Cross-corpora Hate Speech Detection (2022.findings-acl)

Copied to clipboard

Challenge: Hate speech classifiers exhibit performance degradation when evaluated on datasets different from the source.
Approach: They propose to automatically identify and reduce spurious correlations using attribution methods with dynamic refinement of the list of terms that need to be regularized during training.
Outcome: The proposed method improves performance across corpora and on different datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations