The Relevance of Value Systems for Offensive Language Detection (2026.eacl-long)
Copied to clipboard
| Challenge: | Recent research in perspectivism has departed from the assumption that offensiveness can be defined through a universal perspective. |
| Approach: | They propose to use a dataset consisting of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems to identify offensiveness patterns. |
| Outcome: | The proposed dataset consists of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems. |
Similar Papers
Ruddit: Norms of Offensiveness for English Reddit Comments (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods to detect offensive language have been limited by categorical labels . however, there are several challenges in the detection of such content . |
| Approach: | They analyze Reddit comments with fine-grained, real-valued offensiveness scores . they evaluate the ability of widely-used neural models to predict offensiveness . |
| Outcome: | The proposed method produces highly reliable offensiveness scores and can predict scores on reddit comments. |
On the Role of Speech Data in Reducing Toxicity Detection Bias (2025.naacl-long)
Copied to clipboard
Samuel Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà
| Challenge: | Text toxicity detection systems produce disproportionate rates of false positives on demographic groups . toxicity classification systems often misinterpret benign group mentions as toxic . |
| Approach: | They use group annotations to compare text-based and speech-based toxicity detection systems. |
| Outcome: | The results show that access to speech data supports reduced bias against group mentions . the authors recommend improving classifiers, rather than transcription pipelines if possible . |
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing large language models struggle with systematically perturbed data designed to evade detection mechanisms. |
| Approach: | They propose a large language model with homophonic substitutions and emoji transformations to test their models' robustness against cloaking perturbations. |
| Outcome: | The proposed model underperforms in detecting offensive content when perturbations are applied to Chinese language datasets. |
Controversy and Conformity: from Generalized to Personalized Aggressiveness Detection (2021.acl-long)
Copied to clipboard
Kamil Kanclerz, Alicja Figas, Marcin Gruza, Tomasz Kajdanowicz, Jan Kocon, Daria Puchalska, Przemyslaw Kazienko
| Challenge: | a new method to personalize documents that are perceived differently by users is needed . a recent study found that only a few annotations of controversial documents outperform classic methods . |
| Approach: | They propose to use some known, most controversial texts whose offensiveness is very ambiguous . they use user conformity-based measures or embeddings of their previous annotations to improve personalized reasoning . |
| Outcome: | The proposed methods outperform standard methods in document controversy and user nonconformity . the more controversial the content, the greater the gain, the authors say . |
Hate Personified: Investigating the role of LLMs in content moderation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Our work provides preliminary guidelines and highlights the nuances of applying Large Language models in culturally sensitive cases. |
| Approach: | They propose to use large language models to help with content moderation to assess how well the needs of diverse groups are reflected in annotated posts. |
| Outcome: | The proposed model is able to leverage community-based flagging efforts and exposure to adversaries. |
KOAS: Korean Text Offensiveness Analysis System (2021.emnlp-demo)
Copied to clipboard
San-Hee Park, Kang-Min Kim, Seonhee Cho, Jun-Hyung Park, Hyuntae Park, Hyuna Kim, Seongwon Chung, SangKeun Lee
| Challenge: | morphological richness and complex syntax of Korean cause difficulties in neural model training. |
| Approach: | They propose a system that exploits contextual and linguistic features and estimates an offensiveness score for a Korean text. |
| Outcome: | The proposed system exploits both contextual and linguistic features and estimates an offensiveness score for a Korean text. |
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Adapting cultural values in Large Language Models presents significant challenges due to biases and data limitations. |
| Approach: | They propose to augment World Values Survey (WVS) data with encyclopedic and scenario-based cultural narratives from Wikipedia and NormAd to address these limitations. |
| Outcome: | The proposed approach enhances cultural distinctiveness and improves classification performance across cultures. |
Can Language Models Reason about Individualistic Human Values and Preferences? (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods and evaluation frameworks for achieving pluralistic alignment are limited by the diversity of people, which is pre-specified and coarsely categorized, papering over individuality. |
| Approach: | They propose to use a dataset transformed from the influential World Values Survey to study language models on the specific challenge of individualistic value reasoning. |
| Outcome: | The proposed model can predict individualistic values with accuracies between 55% and 65%, while a precise description of individualistic value judgments cannot be approximated only via demographic information. |
An Exploratory Analysis of the Relation between Offensive Language and Mental Health (2021.findings-acl)
Copied to clipboard
| Challenge: | Using computational models, the use of offensive language is pervasive in social media . a popular line of research is the study of machine learning classifiers to identify offensive content online . |
| Approach: | They analyze social media posts written by individuals with depression and those without . they train computational models to compare use of offensive language with depression detection . |
| Outcome: | The proposed models show that offensive language is more frequently used in the samples written by individuals with depression and those showing signs of depression. |
Detecting Community Sensitive Norm Violations in Online Conversations (2021.findings-emnlp)
Copied to clipboard
Chan Young Park, Julia Mendelsohn, Karthik Radhakrishnan, Kinjal Jain, Tushar Kanakagiri, David Jurgens, Yulia Tsvetkov
| Challenge: | Existing efforts to identify unacceptable behavior have focused on toxicity as the sole form of community norm violation. |
| Approach: | They propose a dataset that focuses on a more complete spectrum of community norms and their violations in local conversational and global contexts. |
| Outcome: | The proposed model improves the detection of community norm violations in local conversational and global contexts. |