Abusive language in Spanish children and young teenager’s conversations: data preparation and short text classification with contextual word embeddings (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on how to automatically detect abusive short texts are gaining interest in the natural language processing community. |
| Approach: | They propose to use a contextual word embedding model to automatically detect abusive short texts for Spanish language. |
| Outcome: | The proposed model outperforms classical methods in the detection of abusive short texts for the spanish language. |
Similar Papers
Cross-domain and Cross-lingual Abusive Language Detection: A Hybrid Approach with Deep Learning and a Multilingual Lexicon (P19-2)
Copied to clipboard
| Challenge: | Detecting online abusive language in social media messages is gaining increasing attention from scholars and stakeholders. |
| Approach: | They propose a hybrid approach with deep learning and a multilingual lexicon to cross-domain and cross-lingual detection of abusive content. |
| Outcome: | The proposed system can detect abusive content across domains and languages using a multilingual lexicon and a domain-independent lexical. |
Don’t Augment, Rewrite? Assessing Abusive Language Detection with Synthetic Data (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for abusive language detection and content moderation are limited by regulatory bodies and social media platforms. |
| Approach: | They propose to replace existing datasets in English with synthetic data by rewriting original texts with an instruction-based generative model. |
| Outcome: | The proposed model improves performance in cross-dataset training. |
Humans Need Context, What about Machines? Investigating Conversational Context in Abusive Language Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | In this paper, we examine the role of conversational context in abusive language detection . prior studies have ignored the contextual nature of abusive language, ignoring this aspect . toxicity, hate speech, harmful stereotypes are among the forms of harmful language . |
| Approach: | They propose to use conversational context to analyze abusive language detection using two methods . they use "abusive language" as an umbrella term to refer to various forms of harmful language . |
| Outcome: | The proposed approach is based on two datasets in English and a new dataset of French tweets annotated for hate speech and stereotypes. |
Introducing CAD: the Contextual Abuse Dataset (2021.naacl-main)
Copied to clipboard
| Challenge: | Detecting and classifying online abuse is a complex and nuanced task, despite many advances in the power and availability of computational tools. |
| Approach: | They propose to annotate a reddit conversation thread with six distinct primary and secondary categories and an expert-driven group-adjudication process for high quality annotations. |
| Outcome: | The proposed dataset contains six distinct primary and secondary categories and uses an expert-driven group-adjudication process for high quality annotations. |
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets for abusive language detection are expensive and lack of knowledge about the target is a challenge. |
| Approach: | They propose to build models cheaply for a new target label set and/or language, using only a few training examples of the target domain. |
| Outcome: | The proposed model improves monolingually and across languages using existing datasets and only a few-shots of the target domain. |
Detecting context abusiveness using hierarchical deep learning (D19-50)
Copied to clipboard
| Challenge: | Abusive text is a serious problem in social media and causes many issues among users . a model that detects text abusiveness in context without explicit abusive words is challenging . |
| Approach: | They propose to use an abusive lexicon to determine the existence of an abusive word in text . they combine local and global features to evaluate the model using benchmark data . |
| Outcome: | The proposed model outperforms all previous models for detecting abusiveness in text without abusive words. |
Euphemistic Abuse – A New Dataset and Classification Experiments for Implicitly Abusive Language (2023.emnlp-main)
Copied to clipboard
| Challenge: | Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web. |
| Approach: | They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances. |
| Outcome: | The proposed classifier augments training data with automatically-generated GPT-3 completions. |
Multilingual Offensive Language Identification with Cross-lingual Embeddings (2020.emnlp-main)
Copied to clipboard
| Challenge: | Several studies investigating methods to detect offensive content in social media use English data. |
| Approach: | They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources. |
| Outcome: | The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish. |
Inducing a Lexicon of Abusive Words – a Feature-Based Approach (N18-1)
Copied to clipboard
| Challenge: | a new classification task is needed to identify abusive words among a set of negative polar expressions. |
| Approach: | They propose to calibrate a domain-independent lexicon for detection of abusive words . they use a small manually annotated base lexico to calibrated a large lexical . |
| Outcome: | The proposed feature can be calibrated on a small manually annotated base lexicon and produced on large datasets. |
CyberAgressionAdo-v1: a Dataset of Annotated Online Aggressions in French Collected through a Role-playing Game (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have highlighted that private instant messaging platforms are major mediums of cyber aggression among teens. |
| Approach: | They present a dataset of aggressive chats in French collected through a role-playing game in high-schools . they provide insights on the different types of aggression and verbal abuse depending on the targeted victims . |
| Outcome: | The proposed dataset analyzes aggressive conversations in French on a role-playing game in high schools . it provides insights on the different types of aggression and verbal abuse depending on the targeted victims . |