Challenge: a new framework for analyzing hate speech definitions is proposed to address cultural differences in interpretations . a dataset of 493 definitions from more than 100 cultures is used to analyze hate speech .
Approach: They propose a framework for a cross-cultural and cross-domain analysis of hate speech definitions . they use open-source LLMs to analyze the impact of different definitions on hate speech detection .
Outcome: The proposed framework enables cross-cultural and cross-domain analysis of hate speech definitions . it reveals that many domains borrow definitions from one another without taking into account target culture .

Similar Papers

Multilingual and Multi-Aspect Hate Speech Analysis (D19-1)

Copied to clipboard

Challenge: Current research on hate speech analysis is oriented towards monolingual and single classification tasks.
Approach: They propose to use a multilingual multi-aspect hate speech analysis dataset to test current methods . they evaluate the dataset in various classification settings and discuss how to leverage annotations .
Outcome: The proposed dataset can be used to improve hate speech detection and classification in general.
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis (2024.naacl-long)

Copied to clipboard

Challenge: Existing datasets for hate speech detection neglect the cultural diversity within a single language.
Approach: They propose a CR**oss-cultural **E**nglish **Hate* speech dataset that uses culturally hateful keywords to identify posts from four countries plus the United States.
Outcome: The proposed dataset shows that only 56.2% of the posts in CREHate achieve consensus among all countries, with the highest pairwise label difference rate of 26%.
Explain the Flag: Contextualizing Hate Speech Beyond Censorship (2026.findings-acl)

Copied to clipboard

Challenge: a hybrid approach to detect and explain hate speech combines large language models with vocabularies to detect hate speech in three languages . authors: the spread of hate speech online has serious personal, social, and legal consequences . eu has launched initiatives to analyze, regulate, and counteract online hate speech, authors say .
Approach: They propose a hybrid approach that combines Large Language Models with vocabularies to detect hate speech in English, French, and Greek.
Outcome: The proposed approach outperforms baselines in English, French, and Greek . it uses large language models and vocabularies to detect and explain hate speech . human evaluation shows that the proposed approach is accurate and clear .
Word-Level Detection of Code-Mixed Hate Speech with Multilingual Domain Transfer (2025.findings-acl)

Copied to clipboard

Challenge: a growing problem in language detection tasks is code-mixing, a combination of more than one language . lack of available datasets for code-mixing causes the problem . authors propose a multilingual approach to code-matching .
Approach: They propose to use an annotated hate speech dataset to detect code-mixing in profane language . they propose to apply bilingual fine-tuned models to code-mixed hate speech in german rap lyrics .
Outcome: The proposed model can detect code-mixed hate speech and neologisms in German rap lyrics . the proposed model is more nuanced than binary classification .
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies on explicit or overt hate speech have failed to address a more pervasive form based on coded or indirect language.
Approach: They propose a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication.
Outcome: The proposed dataset will serve as a useful benchmark for understanding this multifaceted issue.
The Challenges of Creating a Parallel Multilingual Hate Speech Corpus: An Exploration (2024.lrec-main)

Copied to clipboard

Challenge: Hate speech is one of the most demanding topics in Natural Language Processing, as its multifaceted nature is accompanied by a handful of challenges, such as multilinguality and cross-linguality.
Approach: They propose a pipeline that could be used to create a parallel multilingual hate speech dataset using machine translation.
Outcome: The proposed pipeline will be able to create a parallel multilingual hate speech dataset using machine translation.
Directions for NLP Practices Applied to Online Hate Speech Detection (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to address hate speech in online spaces have relied on conventions and practices from NLP.
Approach: They argue that many conventions in NLP are poorly suited for the problem and encourage researchers to develop methods that are more appropriate for the task.
Outcome: The proposed methods are poorly suited for the problem and should be adapted to address the propagation of online harms.
Uncovering the Root of Hate Speech: A Dataset for Identifying Hate Instigating Speech (2023.findings-emnlp)

Copied to clipboard

Challenge: a lack of comprehensive datasets specifically annotated for hate instigating speech hinders research . lack of reliable models for hate triggering makes it difficult to apply off-the-shelf models to the problem.
Approach: They propose to use a multilingual dataset to identify hate instigating speech . lack of comprehensive datasets specifically annotated for hate instigators hinders their work .
Outcome: The proposed dataset identifies hate instigating speech across languages . lack of comprehensive datasets makes it difficult to train and evaluate models .
Annotating for Hate Speech: The MaNeCo Corpus and Some Input from Critical Discourse Analysis (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for detecting hate speech are based on the problem of identification, but there is no clear definition of hate speech.
Approach: They propose a multi-layer annotation scheme for the detection of hate speech in a web 2.0 corpus . they propose to use a binary hate speech classification to identify hate speech .
Outcome: The proposed scheme is piloted against a binary hate speech classification and appears to yield higher inter-annotator agreement.
HARE: Explainable Hate Speech Detection with Step-by-Step Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent benchmarks have attempted to identify and explain hate speech but lack the reasoning to supervise detection models.
Approach: They propose a framework that uses large language models to fill in the gaps in hate speech explanations by using existing annotations.
Outcome: The proposed framework outperforms baselines on SBIC and Implicit Hate using model-generated data and improves generalization to unseen datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations