Challenge: generative data augmentation has been shown to be effective in offensive language detection but the potential for bias injection has not been investigated.
Approach: They propose to investigate the robustness of models trained on generated data in a variety of data augmentation setups and analyze models using the HateCheck suite.
Outcome: The proposed model training setups on four English offensive language datasets are robust and robust, while the generative DA setups do not present bias injection issues.

Similar Papers

HateGAN: Adversarial Generative-Based Data Augmentation for Hate Speech Detection (2020.coling-main)

Copied to clipboard

Challenge: Existing methods to detect online hate speech depend heavily on labeled datasets for training, which results in poor detection performance of the hate speech class.
Approach: They propose a deep generative reinforcement learning model which augments two commonly-used hate speech detection datasets with the HateGAN generated tweets.
Outcome: The proposed model improves the detection performance of hate speech class regardless of the classifiers and datasets used in the detection task.
Exploring Data Augmentation Strategies for Hate Speech Detection in Roman Urdu (2022.lrec-1)

Copied to clipboard

Challenge: a number of social media platforms are generating hateful content, a new study finds . augmentation techniques are needed to improve the performance of the models .
Approach: They evaluate different data augmentation techniques for the improvement of hate speech detection in Roman Urdu.
Outcome: The proposed techniques improve hate speech detection in Roman Urdu on two datasets.
Delving into Qualitative Implications of Synthetic Data for Hate Speech Detection (2024.emnlp-main)

Copied to clipboard

Challenge: Recent work on synthetic data for training models for NLP tasks reports mixed results on subjective tasks such as hate speech detection.
Approach: They propose to use synthetic data to train models for highly subjective tasks such as hate speech detection to investigate the potential and specific pitfalls of using it.
Outcome: The proposed model outperforms models trained with real data on hate speech detection tasks, but it fails to accurately reflect real-world data on linguistic dimensions and results in different class distributions.
Don’t Augment, Rewrite? Assessing Abusive Language Detection with Synthetic Data (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for abusive language detection and content moderation are limited by regulatory bodies and social media platforms.
Approach: They propose to replace existing datasets in English with synthetic data by rewriting original texts with an instruction-based generative model.
Outcome: The proposed model improves performance in cross-dataset training.
HARALD: Augmenting Hate Speech Data Sets with Real Data (2022.findings-emnlp)

Copied to clipboard

Challenge: Hate speech detection depends on the availability of variable labeled data.
Approach: They propose a method that uses real unlabelled data from online platforms to augment existing models by harvesting and processing it.
Outcome: The proposed approach improves the classification performance of hate speech classification models.
Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategy (2025.findings-emnlp)

Copied to clipboard

Challenge: Strong attribute control can distort meaning, while prioritizing semantic preservation may weaken attribute alignment.
Approach: They propose a method that restricts accepted samples to text meeting a minimum BERTScore threshold and applies gradient-assisted proposal generation to improve attribute alignment.
Outcome: a new method for counterfactual text generation improves attribute alignment and semantic preservation . the proposed method achieved the best macro F1-score in two of three test sets .
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Hate speech datasets focus on English-language content, hindering effective models . annotating hateful content is expensive, time-consuming and potentially harmful to annotators.
Approach: They propose to use ISO 639-1 codes to fine-tune models on one source language and apply them to another language.
Outcome: The proposed approach performs well on some tasks, but fails on many others.
A little goes a long way: Improving toxic language classification despite data scarcity (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for toxic language classification have not been thoroughly explored.
Approach: They propose to use data augmentation to generate new synthetic data from labeled seed datasets to improve toxic language classification.
Outcome: The proposed techniques perform well on very scarce toxic language datasets while performing worse on shallower models.
Reinforced Counterfactual Data Augmentation for Dual Sentiment Classification (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to improve generalization ability by augmenting training data with synonymous examples or adding random noises to word embeddings cannot address spurious association problem.
Approach: They propose an end-to-end reinforcement learning framework which jointly performs counterfactual data generation and dual sentiment classification.
Outcome: The proposed framework outperforms strong data augmentation baselines on several benchmark sentiment classification datasets.
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection (2023.emnlp-main)

Copied to clipboard

Challenge: Past work has shown that counterfactually augmented data (CADs) can improve models' performance on out-of-domain tests.
Approach: They use Polyjuice, ChatGPT, and Flan-T5 to automatically generate CADs . they find that CAD generates a model that flips the original label with minimal changes .
Outcome: The proposed model improves model robustness on out-of-domain test sets and individual data points.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations