Challenge: a recent study shows that machine learning models are biased and they might make the wrong decisions for the wrong reasons.
Approach: They investigate the impact of social bias on the performance of hate speech detection models . they also investigate the causal effect of intersectional bias on models' unfairness .
Outcome: The proposed model is biased and makes the wrong decisions for the wrong reasons.

Similar Papers

Mitigating Biases in Hate Speech Detection from A Causal Perspective (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to detect hate speech are prone to spurious correlations between training data and labels, which could lead to biased treatment of vulnerable and minority groups.
Approach: They propose to use grammar induction to find grammar patterns for hate speech and analyze this phenomenon from a causal perspective.
Outcome: The proposed methods can detect hate speech from a causal perspective and are robust to different datasets.
The Risk of Racial Bias in Hate Speech Detection (P19-1)

Copied to clipboard

Challenge: Annotators’ insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations.
Approach: They propose *dialect* and *race priming* as ways to reduce the racial bias in hate speech detection models by detecting differences in dialects in annotated tweets.
Outcome: The proposed models acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others.
Robust Hate Speech Detection via Mitigating Spurious Correlations (2022.aacl-short)

Copied to clipboard

Challenge: a novel hate speech detection model can be used to detect word- and character-level adversarial attacks . existing adversarials assume that attackers replace the target words with other names to evade detection .
Approach: They propose a robust hate speech detection model that can defend against adversarial attacks . they describe the process of hate speech recognition by a causal graph and a regularized entropy loss function to quantify spurious correlation .
Outcome: The proposed model can defend against word- and character-level adversarial attacks.
Comparative Evaluation of Label-Agnostic Selection Bias in Multilingual Hate Speech Datasets (2020.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that data collection is neglected by ignoring the quality of data.
Approach: They propose to use latent semantics to evaluate selection bias in hate speech . they compare latent Dirichlet Allocation (LDA) to eleven hate speech corpora .
Outcome: The proposed method could be revisable before focusing on classification performance.
Reducing Gender Bias in Abusive Language Detection (D18-1)

Copied to clipboard

Challenge: Abusive language detection models tend to be biased toward identity words of a certain group of people . recent studies have raised concerns about the robustness of such systems .
Approach: They propose to use debiased word embeddings, gender swap data augmentation to reduce model bias . they also propose to fine-tune models with a larger corpus to correct such bias if needed .
Outcome: The proposed methods reduce model bias by 90-98% and can be extended to correct model bias in other scenarios.
Toxic, Hateful, Offensive or Abusive? What Are We Really Classifying? An Empirical Analysis of Hate Speech Datasets (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that many definitions are being used for equivalent concepts, making most datasets incompatible.
Approach: They analyze six publicly available datasets to determine their similarity and compatibility . they propose to use Fast Text word vectors to analyze similarity between different datasets .
Outcome: The proposed model performs better on similar datasets and worse on more non-offensive samples.
Mind Your Bias: A Critical Review of Bias Detection Methods for Contextual Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for detection of biases in contextual language models are inconsistent and inconclusive.
Approach: They propose to use word embedding association test to detect biases in contextual language models to compare them with other methods.
Outcome: The proposed methods are inconsistent and inconclusive for language models with word embeddings.
Bias and Fairness in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: a tutorial will review the history of bias and fairness studies in machine learning and language processing .
Approach: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models .
Outcome: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks .
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter (2025.acl-long)

Copied to clipboard

Challenge: Prior work on automated hate speech detection models has been limited due to systematic biases in evaluation datasets and poor performance across geographies.
Approach: They propose to construct a global hate speech dataset representative of social media settings from tweets posted on September 21, 2022.
Outcome: The proposed dataset covers eight languages and four English-speaking countries and covers eight countries where English is the main language on Twitter.
Detection of Abusive Language: the Problem of Biased Datasets (N19-1)

Copied to clipboard

Challenge: Recent studies have reported high classification performance on datasets with difficult cases of abusive language.
Approach: They examine the impact of data bias on abusive language detection by focusing on specific microposts rather than random sampling.
Outcome: The proposed method is more accurate and more accurate than random sampling.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations