Academic-Industrial Perspective on the Development and Deployment of a Moderation System for a Newspaper Website (L18-1)
Copied to clipboard
| Challenge: | a system that supports the moderation of user comments on a large newspaper website is described in this paper. |
| Approach: | They describe an approach and experiences from the development, deployment and usability testing of a natural language processing and information retrieval system that supports the moderation of user comments on a large newspaper website. |
| Outcome: | The proposed system supports the moderation of user comments on a large newspaper website. |
Similar Papers
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing knowledge on how and why NLP methods make content moderation decisions is limited . authors examine how and when to use LLMs in content modeation . |
| Approach: | They use Shapley values and LLM-guided explanations to reverse-engineer content moderation decisions across countries. |
| Outcome: | The proposed methods show that they reverse-engineer content moderation decisions across countries and over time. |
Addressing Linguistic Bias through a Contrastive Analysis of Academic Writing in the NLP Domain (2023.emnlp-main)
Copied to clipboard
| Challenge: | a reviewer’s opinion of the nativeness of expression in an academic paper affects the likelihood of it being accepted for publication. |
| Approach: | They conduct a statistical analysis of paper abstracts from the natural language processing domain to identify how authors from different linguistic backgrounds differ in the lexical, morphological, syntactic and cohesive aspects of their writing. |
| Outcome: | The results suggest that there is potential for linguistic bias in the domain of natural language processing. |
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP (2026.acl-long)
Copied to clipboard
Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, Miryam de Lhoneux
| Challenge: | Wikipedia’s perceived high quality and broad language coverage have established it as a fundamental resource in NLP. |
| Approach: | They propose a data filtering procedure which removes a large percentage of Wikipedia's data and a 4-level quality ranking of the site. |
| Outcome: | The results show that the proposed filtering procedure outperforms the raw Wikipedia models in three language modelling scenarios. |
Language (Technology) is Power: A Critical Survey of “Bias” in NLP (2020.acl-main)
Copied to clipboard
| Challenge: | 146 papers analyzing "bias" in NLP systems lack normative reasoning, we find . authors propose three recommendations for work analyzing “bias” in Nlp systems . |
| Approach: | They propose three recommendations for analyzing "bias" in NLP systems . they propose to focus on what kinds of system behaviors are harmful, in what ways, to whom, and why . |
| Outcome: | The proposed methods for measuring or mitigating “bias” are poorly matched to their motivations and do not engage critically with literature outside of NLP. |
Regulation and NLP (RegNLP): Taming Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | polarization in AI safety and ethics debates are swaying political agendas on AI regulation and governance . regulation studies are rich source of knowledge on how to systematically deal with risk and uncertainty . |
| Approach: | They argue that NLP research can benefit from proximity to regulatory studies . they argue that regulation studies should focus on linking scientific knowledge to regulatory processes . |
| Outcome: | The proposed research space should focus on linking scientific knowledge to regulatory processes based on systematic methodologies. |
Multilingual Content Moderation: A Case Study on Reddit (2023.eacl-main)
Copied to clipboard
| Challenge: | a growing need for AI moderators to safeguard users and protect mental health of human moderator from traumatic content. |
| Approach: | They propose to use a multilingual dataset to study the challenges of content moderation . they propose to analyze 1.8 million Reddit comments in English, german, spanish and french . |
| Outcome: | The proposed dataset highlights the challenges and suggests related research problems . it shows that the proposed model can be used to predict the violated rule . |
Synthetic Data for English Lexical Normalization: How Close Can We Get to Manually Annotated Data? (2020.lrec-1)
Copied to clipboard
| Challenge: | Social media data is a valuable data resource for natural language processing tasks. |
| Approach: | They propose to adapt input text to a more standard form, a task also referred to as normalization. |
| Outcome: | The proposed system scores 94.29 accuracy on the test data compared to 95.22 when trained on human-annotated data. |
Internal and External Impacts of Natural Language Processing Papers (2025.acl-short)
Copied to clipboard
| Challenge: | a new study examines the impact of NLP research published in top-tier conferences from 1979 to 2024 . language modeling has the widest internal and external influence, while linguistic foundations have lower impacts . |
| Approach: | They analyze citations from research articles and external sources to determine how NLP topics are consumed internally and externally. |
| Outcome: | The findings show that language modeling has the widest internal and external influence . ethics, bias, and fairness show significant attention in policy documents with fewer academic citations . |
Incorporating Worker Perspectives into MTurk Annotation Practices for NLP (2023.emnlp-main)
Copied to clipboard
| Challenge: | Current approaches to data collection for natural language processing on Amazon Mechanical Turk (MTurk) are susceptible to issues regarding workers’ rights and poor response quality without considering the perspectives of MTurq workers. |
| Approach: | They conducted a critical literature review and a survey of MTurk workers to address open questions regarding fair payment, worker privacy, data quality, and considering worker incentives. |
| Outcome: | The findings suggest that future studies may better account for MTurk workers’ experiences in order to respect workers' rights and improve response quality. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |