Inducing a Lexicon of Abusive Words – a Feature-Based Approach (N18-1)

Copied to clipboard

Challenge: a new classification task is needed to identify abusive words among a set of negative polar expressions.
Approach: They propose to calibrate a domain-independent lexicon for detection of abusive words . they use a small manually annotated base lexico to calibrated a large lexical .
Outcome: The proposed feature can be calibrated on a small manually annotated base lexicon and produced on large datasets.

Similar Papers

Exploiting Emojis for Abusive Language Detection (2021.eacl-main)

Copied to clipboard

Challenge: emojis can be used as a proxy for learning a lexicon of abusive words . eliot safina and samuel khan are the authors of this paper .
Approach: They propose to use abusive emojis as a proxy for learning a lexicon of abusive words.
Outcome: The proposed approach generates a lexicon that performs as well as the most advanced lexical induction method.
Unraveling the Search Space of Abusive Language in Wikipedia with Dynamic Lexicon Acquisition (D19-50)

Copied to clipboard

Challenge: Existing methods to detect abusive language only train one classifier for the whole variety of offending . a new method can support a moderator with explicit unraveled explanations for why something was flagged as abusive .
Approach: a new method is proposed to distinguish explicitly abusive cases from the more "shadowed" ones . the researchers extend a lexicon of abusive terms to include new obfuscations of abusive words .
Outcome: a new method can distinguish explicitly abusive cases from the more "shadowed" ones . the method can support a moderator with explicit unraveled explanations for why something was flagged as abusive .
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets for abusive language detection are expensive and lack of knowledge about the target is a challenge.
Approach: They propose to build models cheaply for a new target label set and/or language, using only a few training examples of the target domain.
Outcome: The proposed model improves monolingually and across languages using existing datasets and only a few-shots of the target domain.
Euphemistic Abuse – A New Dataset and Classification Experiments for Implicitly Abusive Language (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web.
Approach: They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances.
Outcome: The proposed classifier augments training data with automatically-generated GPT-3 completions.
Cross-domain and Cross-lingual Abusive Language Detection: A Hybrid Approach with Deep Learning and a Multilingual Lexicon (P19-2)

Copied to clipboard

Challenge: Detecting online abusive language in social media messages is gaining increasing attention from scholars and stakeholders.
Approach: They propose a hybrid approach with deep learning and a multilingual lexicon to cross-domain and cross-lingual detection of abusive content.
Outcome: The proposed system can detect abusive content across domains and languages using a multilingual lexicon and a domain-independent lexical.
Don’t Augment, Rewrite? Assessing Abusive Language Detection with Synthetic Data (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for abusive language detection and content moderation are limited by regulatory bodies and social media platforms.
Approach: They propose to replace existing datasets in English with synthetic data by rewriting original texts with an instruction-based generative model.
Outcome: The proposed model improves performance in cross-dataset training.
Implicitly Abusive Language – What does it actually look like and why are we not getting there? (2021.naacl-main)

Copied to clipboard

Challenge: Existing datasets make learning implicit abuse difficult, argues a new position paper . a lack of work on implicit abuse has limited the effectiveness of automatic detection .
Approach: They argue that existing datasets make learning implicit abuse difficult . they propose a divide-and-conquer strategy to detect implicit abuse .
Outcome: The proposed model could be improved to detect implicit abuse in a dataset with a standardized model.
Detecting context abusiveness using hierarchical deep learning (D19-50)

Copied to clipboard

Challenge: Abusive text is a serious problem in social media and causes many issues among users . a model that detects text abusiveness in context without explicit abusive words is challenging .
Approach: They propose to use an abusive lexicon to determine the existence of an abusive word in text . they combine local and global features to evaluate the model using benchmark data .
Outcome: The proposed model outperforms all previous models for detecting abusiveness in text without abusive words.
XHate-999: Analyzing and Detecting Abusive Language Across Domains and Languages (2020.coling-main)

Copied to clipboard

Challenge: XHate-999 is a multi-domain and multilingual evaluation data set for abusive language detection . we show that domain- and language-adaption can lead to substantially improved abusive language detecting in the target language .
Approach: They propose a multi-domain and multilingual evaluation data set for abusive language detection that allows for disentanglement of domain transfer and language transfer effects.
Outcome: The proposed model can significantly improve abusive language detection in the target language in the zero-shot transfer setups.
Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation Vectors (2022.acl-long)

Copied to clipboard

Challenge: a new study shows that general abusive language classifiers are reliable in detecting explicit abuse but fail to detect more subtle abuses.
Approach: They propose an interpretability technique to quantify the sensitivity of a trained model to new data . they propose a degree of explicitness metric to suggest out-of-domain unlabeled examples .
Outcome: The proposed interpretability technique is useful for predicting the generalizability of the model on new data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations