ToVo: Toxicity Taxonomy via Voting (2025.findings-naacl)

Copied to clipboard

Challenge: Existing toxic content detection models face limitations due to the closed-source nature of training data and the paucity of explanations for their evaluation mechanism.
Approach: They propose a mechanism that integrates voting and chain-of-thought processes to produce a high-quality open-source dataset for toxic content detection.
Outcome: The proposed model improves transparency and customizability while facilitating better fine-tuning for specific use cases.

Similar Papers

ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models for detecting harmful content lack diversity and quality of datasets.
Approach: They propose a framework for synthesizing toxic information from social media datasets . their framework generates a wide variety of synthetic, yet remarkably realistic, examples of toxic information .
Outcome: The proposed framework can generate a wide variety of synthetic, yet remarkably realistic, examples of toxic information.
From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil Transfer (2022.acl-long)

Copied to clipboard

Challenge: a dataset of English posts with annotations of toxic spans is released . sequence labeling models perform best, but rationale extraction methods are promising .
Approach: They propose a dataset for toxic spans detection that includes an annotation of toxic posts . they propose to add generic rationale extraction mechanisms to the model to obtain toxic span information .
Outcome: The proposed framework is based on a dataset of English posts with toxic span annotations . it shows that sequence labeling models perform best, but that rationale extraction methods are promising .
ModelCitizens: Representing Community Voices in Online Safety (2025.emnlp-main)

Copied to clipboard

Challenge: Existing toxic language detection models are trained on annotations that collapse diverse perspectives into a single ground truth.
Approach: They propose to augment social media posts with conversational scenarios to reflect the impact of conversational context on toxicity.
Outcome: The proposed model outperforms existing models on social media with conversational scenarios.
Detecting Community Sensitive Norm Violations in Online Conversations (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing efforts to identify unacceptable behavior have focused on toxicity as the sole form of community norm violation.
Approach: They propose a dataset that focuses on a more complete spectrum of community norms and their violations in local conversational and global contexts.
Outcome: The proposed model improves the detection of community norm violations in local conversational and global contexts.
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric Method (2024.emnlp-main)

Copied to clipboard

Challenge: Existing efforts to automate content moderation have focused on identifying toxic, offensive, and hateful content . yet, it remains unclear whether improvements have addressed the needs of volunteer content moderators .
Approach: They propose to use a model review to examine the availability of moderators' models to flag violations of various forum rules.
Outcome: The proposed models perform poorly on a significant portion of the rules.
Toxicity, Morality, and Speech Act Guided Stance Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies that focus on stance detection ignore the speech act, toxic, and moral features of tweets or lack an efficient architecture to detect the attitudes across targets.
Approach: They propose a multitasking model that extracts valence, arousal, and dominance aspects hidden in tweets and injects the emotional sense into the embedded text followed by an efficient attention framework to correctly detect the tweet’s stance.
Outcome: The proposed model exploits the toxicity, morality, and speech act features of the tweets to detect the public's stance.
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos (2024.findings-acl)

Copied to clipboard

Challenge: Using a dataset of 931 videos with 4021 code-mixed Hindi-English utterances, we find that video content with multiple modalities is more accurate and more accurate than textual content.
Approach: They propose to use a dataset to analyze toxic content in video content in non-English languages by leveraging language models.
Outcome: The proposed framework achieves an Accuracy and Weighted F1 score of 94.29% and 94.35% for the first time in its class.
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to address toxicity issues with large language models are inadequate . lack of domain-specific knowledge leads to false negatives and excessive sensitivity to toxic speech limits freedom of speech.
Approach: They propose a method that leverages graph search on a meta-toxic knowledge graph to enhance hatred and toxicity detection.
Outcome: The proposed method lowers false positive rate and improves toxicity detection performance in out-of-domain scenarios.
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation (2024.naacl-long)

Copied to clipboard

Challenge: Existing models that focus on explicit toxic speech detection and explanation are prone to error propagation problems . et al., 2018) show that toxic speech models can be prone for generating errors .
Approach: They propose a framework that can detect and explain toxic speech using a target group generator and an encoder-decoder model.
Outcome: The proposed model outperforms baseline models and achieves state-of-the-art effectiveness . the proposed model generates a toxic explanation that matches the ground truth explanation .
TaeBench: Improving Quality of Toxic Adversarial Examples (2025.naacl-industry)

Copied to clipboard

Challenge: Existing adversarial examples generate invalid or ambiguous examples that fool the systems into wrong detection.
Approach: They propose an annotation pipeline for quality control of generated toxic adversarial examples (TAE) they use model-based automated annotation and human-based quality verification to assess quality requirements of a TAE dataset.
Outcome: The proposed pipeline can transfer-attack SOTA toxicity content moderation models and services with adversarial training.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations