Challenge: Existing methods for detecting hate speech are based on the problem of identification, but there is no clear definition of hate speech.
Approach: They propose a multi-layer annotation scheme for the detection of hate speech in a web 2.0 corpus . they propose to use a binary hate speech classification to identify hate speech .
Outcome: The proposed scheme is piloted against a binary hate speech classification and appears to yield higher inter-annotator agreement.

Similar Papers

An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)

Copied to clipboard

Challenge: a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion .
Approach: They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity .
Outcome: The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity.
Untangling Hate Speech Definitions: A Semantic Componential Analysis Across Cultures and Domains (2025.findings-naacl)

Copied to clipboard

Challenge: a new framework for analyzing hate speech definitions is proposed to address cultural differences in interpretations . a dataset of 493 definitions from more than 100 cultures is used to analyze hate speech .
Approach: They propose a framework for a cross-cultural and cross-domain analysis of hate speech definitions . they use open-source LLMs to analyze the impact of different definitions on hate speech detection .
Outcome: The proposed framework enables cross-cultural and cross-domain analysis of hate speech definitions . it reveals that many domains borrow definitions from one another without taking into account target culture .
HateBRXplain: A Benchmark Dataset with Human-Annotated Rationales for Explainable Hate Speech Detection in Brazilian Portuguese (2025.coling-main)

Copied to clipboard

Challenge: Hate speech detection systems have been developed to inhibit offensive and hateful language from being published or spread on the Web and social media.
Approach: They propose to use a Portuguese dataset to provide rationales for hate speech detection with text span annotations.
Outcome: The proposed models outperform the baselines in Portuguese and showed that they provide plausible explanations when compared to human annotations.
HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection (2022.lrec-1)

Copied to clipboard

Challenge: In Brazil, hate speech is prohibited, however the regulation is not effective due to the difficulty of identifying, quantifying and classifying this kind of online content.
Approach: They propose to annotate a large corpus of Brazilian Instagram comments manually and to use it to detect hate speech and offensive language.
Outcome: The HateBR corpus was collected from the comment section of Brazilian politicians’ accounts on Instagram and manually annotated by specialists, reaching a high inter-annotator agreement.
Listening to Affected Communities to Define Extreme Speech: Dataset and Experiments (2022.findings-acl)

Copied to clipboard

Challenge: XTREMESPEECH dataset contains 20,297 social media passages from Brazil, Germany, India and Kenya .
Approach: They propose a hate speech dataset containing 20,297 social media passages from Brazil, Germany, India and Kenya.
Outcome: The proposed dataset contains 20,297 social media passages from Brazil, Germany, India and Kenya.
Directions for NLP Practices Applied to Online Hate Speech Detection (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to address hate speech in online spaces have relied on conventions and practices from NLP.
Approach: They argue that many conventions in NLP are poorly suited for the problem and encourage researchers to develop methods that are more appropriate for the task.
Outcome: The proposed methods are poorly suited for the problem and should be adapted to address the propagation of online harms.
Explain the Flag: Contextualizing Hate Speech Beyond Censorship (2026.findings-acl)

Copied to clipboard

Challenge: a hybrid approach to detect and explain hate speech combines large language models with vocabularies to detect hate speech in three languages . authors: the spread of hate speech online has serious personal, social, and legal consequences . eu has launched initiatives to analyze, regulate, and counteract online hate speech, authors say .
Approach: They propose a hybrid approach that combines Large Language Models with vocabularies to detect hate speech in English, French, and Greek.
Outcome: The proposed approach outperforms baselines in English, French, and Greek . it uses large language models and vocabularies to detect and explain hate speech . human evaluation shows that the proposed approach is accurate and clear .
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies on explicit or overt hate speech have failed to address a more pervasive form based on coded or indirect language.
Approach: They propose a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication.
Outcome: The proposed dataset will serve as a useful benchmark for understanding this multifaceted issue.
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages (2025.naacl-long)

Copied to clipboard

Challenge: Hate speech and abusive language are global phenomena that need sociocultural background knowledge to be understood, identified, and moderated.
Approach: They propose to use a multilingual dataset to collect hate speech and abusive language in 15 African languages to help improve model performance.
Outcome: The proposed datasets are based on tweets annotated by native speakers familiar with the regional culture and show that they perform well in low-resource settings.
So Hateful! Building a Multi-Label Hate Speech Annotated Arabic Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Social media enables widespread propagation of hate speech targeting groups based on ethnicity, religion, or other characteristics.
Approach: They analyze 70,000 Arabic tweets to identify hate speech patterns and train models . 15% of tweets contain offensive language while 6% have hate speech . authors hope to prevent spread of hateful content on social media platforms .
Outcome: The analysis of 70,000 Arabic tweets shows that 15% of tweets contain offensive language while 6% have hate speech . 10% of tweet provide verifiable factual claims, and 7% are deemed important .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations