Challenge: Using a hierarchical tagset, cyberbullying narratives are described in the dataset CyberAgressionAdo-V1 . resulting dataset comprises 19 conversations that have been manually annotated .
Approach: They propose a new tagset that includes tags marking pragmatic-level information occurring in cyberbullying situations.
Outcome: The proposed tagset includes tags marking pragmatic-level information occurring in cyberbullying situations.

Similar Papers

CyberAgressionAdo-v1: a Dataset of Annotated Online Aggressions in French Collected through a Role-playing Game (2022.lrec-1)

Copied to clipboard

Challenge: Recent studies have highlighted that private instant messaging platforms are major mediums of cyber aggression among teens.
Approach: They present a dataset of aggressive chats in French collected through a role-playing game in high-schools . they provide insights on the different types of aggression and verbal abuse depending on the targeted victims .
Outcome: The proposed dataset analyzes aggressive conversations in French on a role-playing game in high schools . it provides insights on the different types of aggression and verbal abuse depending on the targeted victims .
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)

Copied to clipboard

Challenge: 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms.
Approach: They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
Outcome: The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English.
Penetrating Linguistic Disguises: A Slang-aware Label-Aligned Framework for Fine-Grained Toxicity Extraction in Chinese Hate Speech Detection (2026.findings-acl)

Copied to clipboard

Challenge: Flexible word boundaries and linguistic obfuscation, particularly slang, challenge precise span-level hate speech detection in Chinese.
Approach: They propose a Slang-aware Label-Aligned Framework that maps slang to explicit hate semantics and uses task-specific branches to mitigate feature interference.
Outcome: The proposed framework reduces ambiguity by mapping obscure slang to explicit hate semantics.
Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)

Copied to clipboard

Challenge: a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people.
Approach: They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook.
Outcome: The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field.
BullyBench: Youth & Experts-in-the-loop Framework for Intrinsic and Extrinsic Cyberbullying NLP Benchmarking (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing youth-focused CB datasets lack conversational realism and ethical youth involvement with little or no evaluation of their social plausibility.
Approach: They propose a youth-in-the-loop dataset “BullyBench” that incorporates a structured intrinsic quality evaluation with experts-in the-looop (social scientists, psychologists, and content moderators) they perform extrinsic baseline evaluation by benchmarking encoder- and decoder-only language models for multi-class CB role classification.
Outcome: The proposed dataset is evaluated by a team of social scientists, psychologists, and content moderators to assess its quality, relevance, and coherence.
Introducing CAD: the Contextual Abuse Dataset (2021.naacl-main)

Copied to clipboard

Challenge: Detecting and classifying online abuse is a complex and nuanced task, despite many advances in the power and availability of computational tools.
Approach: They propose to annotate a reddit conversation thread with six distinct primary and secondary categories and an expert-driven group-adjudication process for high quality annotations.
Outcome: The proposed dataset contains six distinct primary and secondary categories and uses an expert-driven group-adjudication process for high quality annotations.
Exploring the Emotional Dimension of French Online Toxic Content (2024.lrec-main)

Copied to clipboard

Challenge: Emotion annotations can be used to analyze content and can be applied to content analysis.
Approach: They propose to use a corpus annotation scheme to annotate three online data sets composed of extremist, sexist and hateful messages respectively.
Outcome: The proposed method can provide new insights for content analysis and stronger empirical background for automatic content detection.
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages (2025.naacl-long)

Copied to clipboard

Challenge: Hate speech and abusive language are global phenomena that need sociocultural background knowledge to be understood, identified, and moderated.
Approach: They propose to use a multilingual dataset to collect hate speech and abusive language in 15 African languages to help improve model performance.
Outcome: The proposed datasets are based on tweets annotated by native speakers familiar with the regional culture and show that they perform well in low-resource settings.
HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection (2022.lrec-1)

Copied to clipboard

Challenge: In Brazil, hate speech is prohibited, however the regulation is not effective due to the difficulty of identifying, quantifying and classifying this kind of online content.
Approach: They propose to annotate a large corpus of Brazilian Instagram comments manually and to use it to detect hate speech and offensive language.
Outcome: The HateBR corpus was collected from the comment section of Brazilian politicians’ accounts on Instagram and manually annotated by specialists, reaching a high inter-annotator agreement.
Annotating Online Misogyny (2021.acl-long)

Copied to clipboard

Challenge: Online misogyny is a category of online abusive language with serious and harmful social consequences.
Approach: They propose an iterative annotation process and a taxonomy of labels for annotating misogyny in natural written language and cite a high-quality dataset of annotated posts from social media posts.
Outcome: The proposed method aims to identify misogynistic language in natural written language and annotate it in social media posts using a high-quality dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations