Challenge: Prior research has shown the need to consider community language norms when studying taboo text classification and annotations.
Approach: They propose to use special classifiers tuned for each community's language to study bias in taboo classification and annotation where a community perspective is front and center.
Outcome: The proposed method shows that biases are strongest against African Americans and South Asians . a community perspective is front and center in the proposed method .

Similar Papers

The Risk of Racial Bias in Hate Speech Detection (P19-1)

Copied to clipboard

Challenge: Annotators’ insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations.
Approach: They propose *dialect* and *race priming* as ways to reduce the racial bias in hate speech detection models by detecting differences in dialects in annotated tweets.
Outcome: The proposed models acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others.
Challenges in Automated Debiasing for Toxic Language Detection (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for debiasing toxic language data are limited in their ability to prevent biased behavior in toxic language detection systems.
Approach: They propose to debiase toxic language detection models using lexical and dialectal markers using synthetic labels instead of traditional methods.
Outcome: The proposed method reduces dialectal associations with toxicity despite the use of synthetic labels .
Detecting Community Sensitive Norm Violations in Online Conversations (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing efforts to identify unacceptable behavior have focused on toxicity as the sole form of community norm violation.
Approach: They propose a dataset that focuses on a more complete spectrum of community norms and their violations in local conversational and global contexts.
Outcome: The proposed model improves the detection of community norm violations in local conversational and global contexts.
Toxic, Hateful, Offensive or Abusive? What Are We Really Classifying? An Empirical Analysis of Hate Speech Datasets (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that many definitions are being used for equivalent concepts, making most datasets incompatible.
Approach: They analyze six publicly available datasets to determine their similarity and compatibility . they propose to use Fast Text word vectors to analyze similarity between different datasets .
Outcome: The proposed model performs better on similar datasets and worse on more non-offensive samples.
Social Bias Probing: Fairness Benchmarking for Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for evaluating social biases in language models have been limited to binary association tests on small datasets.
Approach: They propose a framework for probing language models for social biases by assessing disparate treatment . they use a large-scale benchmark to examine the diversity of identities and stereotypes .
Outcome: The proposed framework expands the analysis beyond the binary comparison of stereotypical versus anti-stereotypical identities to include a diverse range of identities and stereotypes.
The “Knowledge–Behavior Gap” in Cultural Taboo Safety of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing cultural benchmarks assess cultural knowledge or values biases, but ignore cultural taboos.
Approach: They propose a benchmark to evaluate and improve the cultural taboo safety of large language models.
Outcome: The proposed benchmark spans 77 countries and regions, and includes over 2,020 taboos.
Towards Identifying Social Bias in Dialog Systems: Framework, Dataset, and Benchmark (2022.findings-emnlp)

Copied to clipboard

Challenge: a number of safety concerns hinder the deployment of open-domain dialog systems, such as offensive languages and toxic behaviors, such social bias is difficult to detect.
Approach: They propose a Dial-Bias Framework for analyzing social bias in conversations . they introduce a Chinese social bias dialog dataset and conduct in-depth ablation studies .
Outcome: The proposed framework is the first annotated Chinese social bias dialog dataset . the proposed framework also provides a fine-grained dialog bias measurement benchmark .
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)

Copied to clipboard

Challenge: 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms.
Approach: They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
Outcome: The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English.
Comparative Evaluation of Label-Agnostic Selection Bias in Multilingual Hate Speech Datasets (2020.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that data collection is neglected by ignoring the quality of data.
Approach: They propose to use latent semantics to evaluate selection bias in hate speech . they compare latent Dirichlet Allocation (LDA) to eleven hate speech corpora .
Outcome: The proposed method could be revisable before focusing on classification performance.
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on human biases are heavily skewed towards Western and European languages . despite growing interest in language models, there are several shortcomings in the literature .
Approach: They scale the Word Embedding Association Test to 24 languages and add culturally relevant information for each language.
Outcome: The proposed language models can reflect and often amplify the effects of bias across linguistic, cultural, and societal borders.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations