| Challenge: | a new classification task is needed to identify abusive words among a set of negative polar expressions. |
| Approach: | They propose to calibrate a domain-independent lexicon for detection of abusive words . they use a small manually annotated base lexico to calibrated a large lexical . |
| Outcome: | The proposed feature can be calibrated on a small manually annotated base lexicon and produced on large datasets. |
Similar Papers
Exploiting Emojis for Abusive Language Detection (2021.eacl-main)
Copied to clipboard
| Challenge: | emojis can be used as a proxy for learning a lexicon of abusive words . eliot safina and samuel khan are the authors of this paper . |
| Approach: | They propose to use abusive emojis as a proxy for learning a lexicon of abusive words. |
| Outcome: | The proposed approach generates a lexicon that performs as well as the most advanced lexical induction method. |
Unraveling the Search Space of Abusive Language in Wikipedia with Dynamic Lexicon Acquisition (D19-50)
Copied to clipboard
| Challenge: | Existing methods to detect abusive language only train one classifier for the whole variety of offending . a new method can support a moderator with explicit unraveled explanations for why something was flagged as abusive . |
| Approach: | a new method is proposed to distinguish explicitly abusive cases from the more "shadowed" ones . the researchers extend a lexicon of abusive terms to include new obfuscations of abusive words . |
| Outcome: | a new method can distinguish explicitly abusive cases from the more "shadowed" ones . the method can support a moderator with explicit unraveled explanations for why something was flagged as abusive . |
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets for abusive language detection are expensive and lack of knowledge about the target is a challenge. |
| Approach: | They propose to build models cheaply for a new target label set and/or language, using only a few training examples of the target domain. |
| Outcome: | The proposed model improves monolingually and across languages using existing datasets and only a few-shots of the target domain. |
Euphemistic Abuse – A New Dataset and Classification Experiments for Implicitly Abusive Language (2023.emnlp-main)
Copied to clipboard
| Challenge: | Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web. |
| Approach: | They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances. |
| Outcome: | The proposed classifier augments training data with automatically-generated GPT-3 completions. |
Cross-domain and Cross-lingual Abusive Language Detection: A Hybrid Approach with Deep Learning and a Multilingual Lexicon (P19-2)
Copied to clipboard
| Challenge: | Detecting online abusive language in social media messages is gaining increasing attention from scholars and stakeholders. |
| Approach: | They propose a hybrid approach with deep learning and a multilingual lexicon to cross-domain and cross-lingual detection of abusive content. |
| Outcome: | The proposed system can detect abusive content across domains and languages using a multilingual lexicon and a domain-independent lexical. |
Don’t Augment, Rewrite? Assessing Abusive Language Detection with Synthetic Data (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for abusive language detection and content moderation are limited by regulatory bodies and social media platforms. |
| Approach: | They propose to replace existing datasets in English with synthetic data by rewriting original texts with an instruction-based generative model. |
| Outcome: | The proposed model improves performance in cross-dataset training. |
Implicitly Abusive Language – What does it actually look like and why are we not getting there? (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing datasets make learning implicit abuse difficult, argues a new position paper . a lack of work on implicit abuse has limited the effectiveness of automatic detection . |
| Approach: | They argue that existing datasets make learning implicit abuse difficult . they propose a divide-and-conquer strategy to detect implicit abuse . |
| Outcome: | The proposed model could be improved to detect implicit abuse in a dataset with a standardized model. |
Detecting context abusiveness using hierarchical deep learning (D19-50)
Copied to clipboard
| Challenge: | Abusive text is a serious problem in social media and causes many issues among users . a model that detects text abusiveness in context without explicit abusive words is challenging . |
| Approach: | They propose to use an abusive lexicon to determine the existence of an abusive word in text . they combine local and global features to evaluate the model using benchmark data . |
| Outcome: | The proposed model outperforms all previous models for detecting abusiveness in text without abusive words. |
XHate-999: Analyzing and Detecting Abusive Language Across Domains and Languages (2020.coling-main)
Copied to clipboard
| Challenge: | XHate-999 is a multi-domain and multilingual evaluation data set for abusive language detection . we show that domain- and language-adaption can lead to substantially improved abusive language detecting in the target language . |
| Approach: | They propose a multi-domain and multilingual evaluation data set for abusive language detection that allows for disentanglement of domain transfer and language transfer effects. |
| Outcome: | The proposed model can significantly improve abusive language detection in the target language in the zero-shot transfer setups. |
Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation Vectors (2022.acl-long)
Copied to clipboard
| Challenge: | a new study shows that general abusive language classifiers are reliable in detecting explicit abuse but fail to detect more subtle abuses. |
| Approach: | They propose an interpretability technique to quantify the sensitivity of a trained model to new data . they propose a degree of explicitness metric to suggest out-of-domain unlabeled examples . |
| Outcome: | The proposed interpretability technique is useful for predicting the generalizability of the model on new data. |