Challenge: Currently, there are three main branches of violence detection, including surveillance of potential threats in offline situation and automatic prevention of harmful media.
Approach: They propose to use the Korean Crime Dialogue Dataset to classify violence that occurs in offline scenarios.
Outcome: The proposed dataset shows that understanding varying relationships among interlocutors improves the performance of crime dialogue classification.

Similar Papers

KOLD: Korean Offensive Language Dataset (2022.emnlp-main)

Copied to clipboard

Challenge: Recent directions for offensive language detection focus on English and do not transfer well to other languages because of cultural and linguistic differences.
Approach: They present a Korean offensive language dataset annotated with offensive language comments . they use the comments as training data for Korean BERT and RoBERTa models .
Outcome: The proposed model improves offensiveness detection, target classification, and span detection while having room for improvement for target group classification and span prediction.
K-HATERS: A Hate Speech Detection Corpus in Korean with Target-Specific Ratings (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets on hate speech detection focus on overt forms of hate . however, a majority of these resources are English-centric, focusing on overtones of hate.
Approach: They propose a new corpus for hate speech detection in Korean with target-specific offensiveness ratings that offer a three-point Likert scale.
Outcome: The proposed corpus is the largest offensive language corpus in Korean and offers target-specific ratings on a three-point Likert scale.
K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment (2022.coling-1)

Copied to clipboard

Challenge: Online hate speech detection resources in other languages are limited.
Approach: They introduce a new dataset for hate speech detection that handles Korean language patterns.
Outcome: The proposed dataset outperforms existing datasets in Korean language classifications.
A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech (2024.acl-long)

Copied to clipboard

Challenge: Using data from 420k Twitter posts, we characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection.
Approach: They develop a codebook to characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection.
Outcome: The proposed codebook analyzes 420k tweets over 3 years and compares classifiers with hateful speech classifier classifier to detect hateful content.
KOAS: Korean Text Offensiveness Analysis System (2021.emnlp-demo)

Copied to clipboard

Challenge: morphological richness and complex syntax of Korean cause difficulties in neural model training.
Approach: They propose a system that exploits contextual and linguistic features and estimates an offensiveness score for a Korean text.
Outcome: The proposed system exploits both contextual and linguistic features and estimates an offensiveness score for a Korean text.
MPDD: A Multi-Party Dialogue Dataset for Analysis of Emotions and Interpersonal Relationships (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets with emotion and relation labels for dialogues are limited.
Approach: They use a Chinese dialogue dataset to annotate emotions and interpersonal relationships on each utterance.
Outcome: The proposed dataset contains 25,548 utterances from 4,142 dialogues.
Introducing CAD: the Contextual Abuse Dataset (2021.naacl-main)

Copied to clipboard

Challenge: Detecting and classifying online abuse is a complex and nuanced task, despite many advances in the power and availability of computational tools.
Approach: They propose to annotate a reddit conversation thread with six distinct primary and secondary categories and an expert-driven group-adjudication process for high quality annotations.
Outcome: The proposed dataset contains six distinct primary and secondary categories and uses an expert-driven group-adjudication process for high quality annotations.
KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained Counselors (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have explored using large language models to augment counseling dialogue datasets, but data from real-world counseling environments may suffer from limited diversity and authenticity.
Approach: They propose to use a Japanese psychological counseling dialogue dataset to simulate counselor-client interactions by using open-source LLMs.
Outcome: The proposed model improves the quality of generated counseling responses and the automatic evaluation of counseling dialogues.
KoCoSa: Korean Context-aware Sarcasm Detection Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Sarcasm is a form of verbal irony where someone says the opposite of what they mean . misunderstanding this sarcasm may lead to fatal errors in dialogue systems .
Approach: They propose a dataset for the Korean dialogue sarcasm detection task that uses 12.8K daily Korean dialogues and the labels on the last response.
Outcome: The proposed system outperforms strong baselines like large language models in the Korean sarcasm detection task.
APEACH: Attacking Pejorative Expressions with Analysis on Crowd-Generated Hate Speech Evaluation Datasets (2022.findings-emnlp)

Copied to clipboard

Challenge: flaming or trolling in online communities is considered hostile behavior . a dataset of hate speech examples can be useful for detecting toxic or pejorative expressions . annotating on existing web text has several limitations that deter the dataset's reliability .
Approach: They propose a dataset that asks users to generate hate speech examples followed by minimal post-labeling.
Outcome: a new approach can collect useful datasets that are less sensitive to overlaps, the authors say . annotating on web text has several limitations that deter the dataset's reliability .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations