Challenge: Hate speech classifiers do not perform equally well in detecting hateful expressions towards different target identities.
Approach: They propose to use two recently proposed functionality test datasets to analyze the impact of different factors on HS prediction.
Outcome: The proposed classifiers do not perform equally well across different datasets and different target identities.

Similar Papers

Latent Hatred: A Benchmark for Understanding Implicit Hate Speech (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies on explicit or overt hate speech have failed to address a more pervasive form based on coded or indirect language.
Approach: They propose a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication.
Outcome: The proposed dataset will serve as a useful benchmark for understanding this multifaceted issue.
Comparative Evaluation of Label-Agnostic Selection Bias in Multilingual Hate Speech Datasets (2020.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that data collection is neglected by ignoring the quality of data.
Approach: They propose to use latent semantics to evaluate selection bias in hate speech . they compare latent Dirichlet Allocation (LDA) to eleven hate speech corpora .
Outcome: The proposed method could be revisable before focusing on classification performance.
The Risk of Racial Bias in Hate Speech Detection (P19-1)

Copied to clipboard

Challenge: Annotators’ insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations.
Approach: They propose *dialect* and *race priming* as ways to reduce the racial bias in hate speech detection models by detecting differences in dialects in annotated tweets.
Outcome: The proposed models acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others.
SOS: Systematic Offensive Stereotyping Bias in Word Embeddings (2022.coling-1)

Copied to clipboard

Challenge: Systematic Offensive Stereotyping (SOS) in word embeddings could lead to associating marginalised groups with hate speech and profanity.
Approach: They propose a quantitative measure of the systematic offensive stereotyping (SOS) in word embeddings and validate it in most commonly used word embeds.
Outcome: The proposed measure correlates with published statistics on online extremism, but does not explain hate speech detection models.
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter (2025.acl-long)

Copied to clipboard

Challenge: Prior work on automated hate speech detection models has been limited due to systematic biases in evaluation datasets and poor performance across geographies.
Approach: They propose to construct a global hate speech dataset representative of social media settings from tweets posted on September 21, 2022.
Outcome: The proposed dataset covers eight languages and four English-speaking countries and covers eight countries where English is the main language on Twitter.
Mitigating Biases in Hate Speech Detection from A Causal Perspective (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to detect hate speech are prone to spurious correlations between training data and labels, which could lead to biased treatment of vulnerable and minority groups.
Approach: They propose to use grammar induction to find grammar patterns for hate speech and analyze this phenomenon from a causal perspective.
Outcome: The proposed methods can detect hate speech from a causal perspective and are robust to different datasets.
Unmasking the Hidden Meaning: Bridging Implicit and Explicit Hate Speech Embedding Representations (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to detect explicit hate speech (HS) are focusing on detecting explicit forms of hateful expressions on user-generated content.
Approach: They propose to examine the differences between embedding implicit and explicit hateful messages . they compare and link explicit and implicit hateful message across datasets .
Outcome: The proposed model improves on explicit hate speech detection while retaining high performance on borderline cases.
Toxic, Hateful, Offensive or Abusive? What Are We Really Classifying? An Empirical Analysis of Hate Speech Datasets (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that many definitions are being used for equivalent concepts, making most datasets incompatible.
Approach: They analyze six publicly available datasets to determine their similarity and compatibility . they propose to use Fast Text word vectors to analyze similarity between different datasets .
Outcome: The proposed model performs better on similar datasets and worse on more non-offensive samples.
Hate Speech Classifiers Learn Normative Social Stereotypes (2023.tacl-1)

Copied to clipboard

Challenge: Social stereotypes negatively impact individuals’ judgments about different groups and may have a critical role in understanding language directed toward marginalized groups.
Approach: They first investigate the impact of novice annotators’ stereotypes on their hate-speech-annotation behavior. Then, they examine the effect of normative stereotypes in language on the aggregated annotated judgments.
Outcome: The framework provides insights into sources of bias in hate-speech moderation, informing ongoing debates regarding machine learning fairness.
Who Speaks Matters: Analysing the Influence of the Speaker’s Linguistic Identity on Hate Classification (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models are known to be brittle and biased against marginalised communities and dialects.
Approach: They investigate the robustness of hate speech classification using LLMs when explicit and implicit markers of the speaker’s ethnicity are injected into the input.
Outcome: The proposed model is robust when explicit and implicit markers of speaker's ethnicity are injected into the input.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations