Challenge: Existing efforts to automate content moderation have focused on identifying toxic, offensive, and hateful content . yet, it remains unclear whether improvements have addressed the needs of volunteer content moderators .
Approach: They propose to use a model review to examine the availability of moderators' models to flag violations of various forum rules.
Outcome: The proposed models perform poorly on a significant portion of the rules.

Similar Papers

Detecting Community Sensitive Norm Violations in Online Conversations (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing efforts to identify unacceptable behavior have focused on toxicity as the sole form of community norm violation.
Approach: They propose a dataset that focuses on a more complete spectrum of community norms and their violations in local conversational and global contexts.
Outcome: The proposed model improves the detection of community norm violations in local conversational and global contexts.
ModelCitizens: Representing Community Voices in Online Safety (2025.emnlp-main)

Copied to clipboard

Challenge: Existing toxic language detection models are trained on annotations that collapse diverse perspectives into a single ground truth.
Approach: They propose to augment social media posts with conversational scenarios to reflect the impact of conversational context on toxicity.
Outcome: The proposed model outperforms existing models on social media with conversational scenarios.
It Is Not Only the Negative that Deserves Attention! Understanding, Generation & Evaluation of (Positive) Moderation (2025.naacl-long)

Copied to clipboard

Challenge: Moderation is essential for maintaining and improving the quality of online discussions.
Approach: They annotate a dataset on 13 modes of discussion and use it to generate positive moderation.
Outcome: The proposed model shows that professional moderation generates higher ratings than professional moderated moderation, but prefers professional moderate in pairwise comparison.
ToVo: Toxicity Taxonomy via Voting (2025.findings-naacl)

Copied to clipboard

Challenge: Existing toxic content detection models face limitations due to the closed-source nature of training data and the paucity of explanations for their evaluation mechanism.
Approach: They propose a mechanism that integrates voting and chain-of-thought processes to produce a high-quality open-source dataset for toxic content detection.
Outcome: The proposed model improves transparency and customizability while facilitating better fine-tuning for specific use cases.
Model-Dependent Moderation: Inconsistencies in Hate Speech Detection Across LLM-based Systems (2025.findings-acl)

Copied to clipboard

Challenge: Content moderation systems powered by large language models are increasingly deployed to detect hate speech . if two systems produce different outcomes for the same content, it undermines consistency and predictability .
Approach: They analyze 1.3+ million sentences from a factorial design to determine hate speech classification . they find identical content receives markedly different classification values across systems .
Outcome: The proposed model finds that identical content receives markedly different classification values across systems.
SLM-Mod: Small Language Models Surpass LLMs at Content Moderation (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are expensive to query in real-time and do not allow for a community-specific approach to content moderation.
Approach: They propose to use small language models for community-specific content moderation tasks by fine-tuning and evaluating their performance against larger open- and closed-sourced models.
Outcome: The proposed models outperform zero-shot LLMs in content moderation tasks with 11.5% higher accuracy and 25.7% higher recall across all communities.
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to train classifiers that predict norm violations are often opacity-prone . a new approach to identify and extract these implicit criteria from historical moderation data is proposed .
Approach: They propose to extract implicit criteria from historical moderation data using an interpretable architecture.
Outcome: The proposed model replicates neural moderation models while providing transparent insights into decision-making processes.
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media (2026.acl-long)

Copied to clipboard

Challenge: Social media are shifting towards community-governed platforms where groups define their own norms.
Approach: They propose a multimodal, multilingual benchmark for detecting 13,371 rule violations across 1,989 Reddit communities . they show that bigger models and increased context provide marginal gains, and universal rules like civility and self-promotion are easier to detect.
Outcome: The proposed model can detect 13,371 rule violations across 1,989 Reddit communities across 2,885 rules in 9 languages.
Moderation in the Wild: Investigating User-Driven Moderation in Online Discussions (2024.eacl-long)

Copied to clipboard

Challenge: Effective content moderation is imperative for fostering healthy and productive discussions in online domains.
Approach: They propose to document and release a dataset of comments in which users act as moderators.
Outcome: The proposed dataset contains 1000 comment-reply pairs with crowdsourced annotations from a large annotator pool and fine-grained annotation schema targeting the functions of moderation, stylistic properties(aggressiveness, subjectivity, sentiment), constructiveness, and individual perspectives of the annotators on the task.
Can Language Model Moderators Improve the Health of Online Discourse? (2024.naacl-long)

Copied to clipboard

Challenge: Existing efforts to automate conversational moderation have focused on banning harmful comments or deleting them, but such efforts can inadvertently push users towards echo chambers that exacerbate polarization.
Approach: They propose a framework to assess models’ moderation capabilities independently of human intervention and propose 'conversational moderation' they propose to use language models as conversational moderators to provide specific feedback on toxic behavior but struggle to influence users to increase their levels of respect and cooperation.
Outcome: The proposed framework assesses models’ moderation capabilities independently of human intervention and shows that appropriately prompted models provide specific and fair feedback on toxic behavior but struggle to influence users to increase their levels of respect and cooperation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations