Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse (2025.acl-long)
Copied to clipboard
| Challenge: | specialized Polish language models are more effective at detecting harmful content than traditional methods. |
| Approach: | They propose a Polish-language dataset for erotic content detection that captures ambiguity, violence, and socially unacceptable behaviors. |
| Outcome: | The proposed dataset shows that specialized Polish language models achieve superior performance compared to multilingual alternatives, with transformer-based architectures showing particular strength in handling imbalanced categories. |
Similar Papers
BAN-PL: A Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl Web Service (2024.lrec-main)
Copied to clipboard
Anna Kolos, Inez Okulska, Kinga Głąbińska, Agnieszka Karlinska, Emilia Wisnios, Paweł Ellerik, Andrzej Prałat
| Challenge: | a new dataset of offensive social media content for the Polish language is presented to address this gap . access to accurate and non-synthetic datasets of social media is limited for low-resource languages . |
| Approach: | They present a new open dataset of offensive social media content for the Polish language . authors propose to make the dataset publicly available to improve access . |
| Outcome: | The proposed dataset includes 691,662 posts and comments from the Polish Reddit . the authors describe the dataset and apply it to real-life content moderation processes . |
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection (2025.coling-main)
Copied to clipboard
| Challenge: | Existing datasets and models fail to address the complexities of multilingual data, authors say . detection of radical content on online platforms has become an increasingly pressing concern . |
| Approach: | They propose a publicly available multilingual dataset annotated with radicalization levels, calls for action, and named entities in English, French, and Arabic. |
| Outcome: | The proposed dataset is annotated with radicalization levels, calls for action, and named entities in English, French, and Arabic. |
Introducing CAD: the Contextual Abuse Dataset (2021.naacl-main)
Copied to clipboard
| Challenge: | Detecting and classifying online abuse is a complex and nuanced task, despite many advances in the power and availability of computational tools. |
| Approach: | They propose to annotate a reddit conversation thread with six distinct primary and secondary categories and an expert-driven group-adjudication process for high quality annotations. |
| Outcome: | The proposed dataset contains six distinct primary and secondary categories and uses an expert-driven group-adjudication process for high quality annotations. |
CoRAL: a Context-aware Croatian Abusive Language Dataset (2022.findings-aacl)
Copied to clipboard
| Challenge: | Semi-automated comment moderation systems can greatly aid human moderators by either automatically classifying the examples or allowing the moderator to prioritize which comments to consider first. |
| Approach: | They propose to use a language and culturally aware Croatian Abusive dataset to analyze inappropriate comments in a context-based manner. |
| Outcome: | The proposed dataset shows that current models degrade when comments are not explicit and further degrades when language skill and context knowledge are required to interpret the comment. |
Multilingual Content Moderation: A Case Study on Reddit (2023.eacl-main)
Copied to clipboard
| Challenge: | a growing need for AI moderators to safeguard users and protect mental health of human moderator from traumatic content. |
| Approach: | They propose to use a multilingual dataset to study the challenges of content moderation . they propose to analyze 1.8 million Reddit comments in English, german, spanish and french . |
| Outcome: | The proposed dataset highlights the challenges and suggests related research problems . it shows that the proposed model can be used to predict the violated rule . |
He said “who’s gonna take care of your children when you are at ACL?”: Reported Sexist Acts are Not Sexist (2020.acl-main)
Copied to clipboard
Patricia Chiril, Véronique Moriceau, Farah Benamara, Alda Mari, Gloria Origgi, Marlène Coulomb-Gully
| Challenge: | Sexism is prejudice or discrimination based on a person's gender. |
| Approach: | They propose to use a French dataset annotated for sexism detection to characterize sexist content and to train deep learning experiments on tweets. |
| Outcome: | The proposed dataset is the first to be used for sexism detection in France and constitutes a first step towards offensive content moderation. |
Evaluation of Sentence Representations in Polish (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for learning sentence representations have been limited in low-resource languages such as Polish . |
| Approach: | They propose two new Polish datasets for evaluating sentence embeddings and evaluate eight different methods including Polish and multilingual models. |
| Outcome: | The proposed methods show strengths and weaknesses in Polish and multilingual models. |
CONAN - COunter NArratives through Nichesourcing: a Multilingual Dataset of Responses to Fight Online Hate Speech (P19-1)
Copied to clipboard
| Challenge: | Davidson et al., 2017): social media platforms and governmental organizations have taken steps to tackle hate speech . Davidson and Norton, 2017: a dataset of hate-speech/counter-narrative pairs is created . authors: identifying hate speech is challenging for the broadness and nuances in cultures and languages . |
| Approach: | They propose to build a large-scale, multilingual, expert-based dataset of hate-speech/counter-narrative pairs . they provide additional annotations about expert demographics, hate and response type . |
| Outcome: | The proposed dataset provides an analysis of hate-speech/counter-narrative pairs in three languages. |
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets for abusive language detection are expensive and lack of knowledge about the target is a challenge. |
| Approach: | They propose to build models cheaply for a new target label set and/or language, using only a few training examples of the target domain. |
| Outcome: | The proposed model improves monolingually and across languages using existing datasets and only a few-shots of the target domain. |
CyberAgressionAdo-v1: a Dataset of Annotated Online Aggressions in French Collected through a Role-playing Game (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have highlighted that private instant messaging platforms are major mediums of cyber aggression among teens. |
| Approach: | They present a dataset of aggressive chats in French collected through a role-playing game in high-schools . they provide insights on the different types of aggression and verbal abuse depending on the targeted victims . |
| Outcome: | The proposed dataset analyzes aggressive conversations in French on a role-playing game in high schools . it provides insights on the different types of aggression and verbal abuse depending on the targeted victims . |