Challenge: Existing datasets on hate speech detection focus on overt forms of hate . however, a majority of these resources are English-centric, focusing on overtones of hate.
Approach: They propose a new corpus for hate speech detection in Korean with target-specific offensiveness ratings that offer a three-point Likert scale.
Outcome: The proposed corpus is the largest offensive language corpus in Korean and offers target-specific ratings on a three-point Likert scale.

Similar Papers

K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment (2022.coling-1)

Copied to clipboard

Challenge: Online hate speech detection resources in other languages are limited.
Approach: They introduce a new dataset for hate speech detection that handles Korean language patterns.
Outcome: The proposed dataset outperforms existing datasets in Korean language classifications.
KOLD: Korean Offensive Language Dataset (2022.emnlp-main)

Copied to clipboard

Challenge: Recent directions for offensive language detection focus on English and do not transfer well to other languages because of cultural and linguistic differences.
Approach: They present a Korean offensive language dataset annotated with offensive language comments . they use the comments as training data for Korean BERT and RoBERTa models .
Outcome: The proposed model improves offensiveness detection, target classification, and span detection while having room for improvement for target group classification and span prediction.
KOAS: Korean Text Offensiveness Analysis System (2021.emnlp-demo)

Copied to clipboard

Challenge: morphological richness and complex syntax of Korean cause difficulties in neural model training.
Approach: They propose a system that exploits contextual and linguistic features and estimates an offensiveness score for a Korean text.
Outcome: The proposed system exploits both contextual and linguistic features and estimates an offensiveness score for a Korean text.
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies on explicit or overt hate speech have failed to address a more pervasive form based on coded or indirect language.
Approach: They propose a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication.
Outcome: The proposed dataset will serve as a useful benchmark for understanding this multifaceted issue.
Uncovering the Root of Hate Speech: A Dataset for Identifying Hate Instigating Speech (2023.findings-emnlp)

Copied to clipboard

Challenge: a lack of comprehensive datasets specifically annotated for hate instigating speech hinders research . lack of reliable models for hate triggering makes it difficult to apply off-the-shelf models to the problem.
Approach: They propose to use a multilingual dataset to identify hate instigating speech . lack of comprehensive datasets specifically annotated for hate instigators hinders their work .
Outcome: The proposed dataset identifies hate instigating speech across languages . lack of comprehensive datasets makes it difficult to train and evaluate models .
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Hate speech datasets focus on English-language content, hindering effective models . annotating hateful content is expensive, time-consuming and potentially harmful to annotators.
Approach: They propose to use ISO 639-1 codes to fine-tune models on one source language and apply them to another language.
Outcome: The proposed approach performs well on some tasks, but fails on many others.
Offensive Language and Hate Speech Detection for Danish (2020.lrec-1)

Copied to clipboard

Challenge: a growing number of social media platforms are detecting and dealing with offensive language . a recent study found that the best performing system for English is best for Danish .
Approach: They propose automatic methods to detect offensive language on social media platforms . they use user-generated comments from various social media sites to find offensive language .
Outcome: The proposed system performs best for both English and Danish language . it achieves a macro averaged F1-score of 0.74 and a best for Danish achieves 0.73 .
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter (2025.acl-long)

Copied to clipboard

Challenge: Prior work on automated hate speech detection models has been limited due to systematic biases in evaluation datasets and poor performance across geographies.
Approach: They propose to construct a global hate speech dataset representative of social media settings from tweets posted on September 21, 2022.
Outcome: The proposed dataset covers eight languages and four English-speaking countries and covers eight countries where English is the main language on Twitter.
What the #?*!: Disentangling Hate Across Target Identities (2025.naacl-long)

Copied to clipboard

Challenge: Hate speech classifiers do not perform equally well in detecting hateful expressions towards different target identities.
Approach: They propose to use two recently proposed functionality test datasets to analyze the impact of different factors on HS prediction.
Outcome: The proposed classifiers do not perform equally well across different datasets and different target identities.
APEACH: Attacking Pejorative Expressions with Analysis on Crowd-Generated Hate Speech Evaluation Datasets (2022.findings-emnlp)

Copied to clipboard

Challenge: flaming or trolling in online communities is considered hostile behavior . a dataset of hate speech examples can be useful for detecting toxic or pejorative expressions . annotating on existing web text has several limitations that deter the dataset's reliability .
Approach: They propose a dataset that asks users to generate hate speech examples followed by minimal post-labeling.
Outcome: a new approach can collect useful datasets that are less sensitive to overlaps, the authors say . annotating on web text has several limitations that deter the dataset's reliability .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations