Moderation in the Wild: Investigating User-Driven Moderation in Online Discussions (2024.eacl-long)
Copied to clipboard
| Challenge: | Effective content moderation is imperative for fostering healthy and productive discussions in online domains. |
| Approach: | They propose to document and release a dataset of comments in which users act as moderators. |
| Outcome: | The proposed dataset contains 1000 comment-reply pairs with crowdsourced annotations from a large annotator pool and fine-grained annotation schema targeting the functions of moderation, stylistic properties(aggressiveness, subjectivity, sentiment), constructiveness, and individual perspectives of the annotators on the task. |
Similar Papers
PerspectiveMod: A Perspectivist Resource for Deliberative Moderation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Human moderators in online discussions face a heterogeneous range of tasks that go beyond content moderation, or policing. |
| Approach: | They propose a dataset of online comments annotated for the question "Does this comment require moderation?" they aim to improve discussion quality by analyzing annotator perspectives and annotating their views. |
| Outcome: | The proposed model is unique in its intentional variation across the level of moderation experience embedded in the source data, the annotator profiles and the individuality of the annnotator. |
It Is Not Only the Negative that Deserves Attention! Understanding, Generation & Evaluation of (Positive) Moderation (2025.naacl-long)
Copied to clipboard
| Challenge: | Moderation is essential for maintaining and improving the quality of online discussions. |
| Approach: | They annotate a dataset on 13 modes of discussion and use it to generate positive moderation. |
| Outcome: | The proposed model shows that professional moderation generates higher ratings than professional moderated moderation, but prefers professional moderate in pairwise comparison. |
Can Language Model Moderators Improve the Health of Online Discourse? (2024.naacl-long)
Copied to clipboard
Hyundong Cho, Shuai Liu, Taiwei Shi, Darpan Jain, Basem Rizk, Yuyang Huang, Zixun Lu, Nuan Wen, Jonathan Gratch, Emilio Ferrara, Jonathan May
| Challenge: | Existing efforts to automate conversational moderation have focused on banning harmful comments or deleting them, but such efforts can inadvertently push users towards echo chambers that exacerbate polarization. |
| Approach: | They propose a framework to assess models’ moderation capabilities independently of human intervention and propose 'conversational moderation' they propose to use language models as conversational moderators to provide specific feedback on toxic behavior but struggle to influence users to increase their levels of respect and cooperation. |
| Outcome: | The proposed framework assesses models’ moderation capabilities independently of human intervention and shows that appropriately prompted models provide specific and fair feedback on toxic behavior but struggle to influence users to increase their levels of respect and cooperation. |
REM: Efficient Semi-Automated Real-Time Moderation of Online Forums (2021.acl-demo)
Copied to clipboard
| Challenge: | REM is a tool for the semi-automated real-time moderation of large scale online forums. |
| Approach: | They propose a semi-automated real-time moderation tool for large scale online forums that maximizes the efficiency of manual moderation by targeting only those comments for which human intervention is needed. |
| Outcome: | The proposed method maximizes the efficiency of manual moderation by targeting only those comments for which human intervention is needed, e.g. due to high classification uncertainty. |
Multilingual Content Moderation: A Case Study on Reddit (2023.eacl-main)
Copied to clipboard
| Challenge: | a growing need for AI moderators to safeguard users and protect mental health of human moderator from traumatic content. |
| Approach: | They propose to use a multilingual dataset to study the challenges of content moderation . they propose to analyze 1.8 million Reddit comments in English, german, spanish and french . |
| Outcome: | The proposed dataset highlights the challenges and suggests related research problems . it shows that the proposed model can be used to predict the violated rule . |
WHoW: A Cross-domain Approach for Analysing Conversation Moderation (2025.naacl-long)
Copied to clipboard
| Challenge: | Using this framework, we annotated 5,657 sentences with human judges and 15,494 sentences with GPT-4o from two domains: TV debates and radio panel discussions. |
| Approach: | They propose an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who). |
| Outcome: | The framework is generalisable across domains and reveals distinct modes of moderation: debate moderators emphasise coordination and facilitate interaction through questions and instructions, panel discussion moderator prioritize information provision and actively participate in discussions. |
A System for Dynamically Tracking Content Moderation on Reddit (2026.acl-demo)
Copied to clipboard
| Challenge: | Recent work in social media platforms delegate content moderation decisions to users and communities. |
| Approach: | They propose a software system for the dynamic monitoring of Reddit posts, communities, and moderation actions to enable scalable and reproducible research on decentralized platform governance and content moderation. |
| Outcome: | The proposed system is the only available solution for general-purpose, real-time, policy-compliant longitudinal data collection on Reddit. |
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric Method (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing efforts to automate content moderation have focused on identifying toxic, offensive, and hateful content . yet, it remains unclear whether improvements have addressed the needs of volunteer content moderators . |
| Approach: | They propose to use a model review to examine the availability of moderators' models to flag violations of various forum rules. |
| Outcome: | The proposed models perform poorly on a significant portion of the rules. |
MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches for content moderation require a separate model for every community and are opaque in their decision-making. |
| Approach: | They propose a modular framework that adds post-hoc explanations to enable scalable content moderation. |
| Outcome: | The proposed framework yields scalable, transparent moderation without fine-tuning across domains. |
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to train classifiers that predict norm violations are often opacity-prone . a new approach to identify and extract these implicit criteria from historical moderation data is proposed . |
| Approach: | They propose to extract implicit criteria from historical moderation data using an interpretable architecture. |
| Outcome: | The proposed model replicates neural moderation models while providing transparent insights into decision-making processes. |