Challenge: Recent research in perspectivism has departed from the assumption that offensiveness can be defined through a universal perspective.
Approach: They propose to use a dataset consisting of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems to identify offensiveness patterns.
Outcome: The proposed dataset consists of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems.

Similar Papers

Ruddit: Norms of Offensiveness for English Reddit Comments (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to detect offensive language have been limited by categorical labels . however, there are several challenges in the detection of such content .
Approach: They analyze Reddit comments with fine-grained, real-valued offensiveness scores . they evaluate the ability of widely-used neural models to predict offensiveness .
Outcome: The proposed method produces highly reliable offensiveness scores and can predict scores on reddit comments.
On the Role of Speech Data in Reducing Toxicity Detection Bias (2025.naacl-long)

Copied to clipboard

Challenge: Text toxicity detection systems produce disproportionate rates of false positives on demographic groups . toxicity classification systems often misinterpret benign group mentions as toxic .
Approach: They use group annotations to compare text-based and speech-based toxicity detection systems.
Outcome: The results show that access to speech data supports reduced bias against group mentions . the authors recommend improving classifiers, rather than transcription pipelines if possible .
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing large language models struggle with systematically perturbed data designed to evade detection mechanisms.
Approach: They propose a large language model with homophonic substitutions and emoji transformations to test their models' robustness against cloaking perturbations.
Outcome: The proposed model underperforms in detecting offensive content when perturbations are applied to Chinese language datasets.
Controversy and Conformity: from Generalized to Personalized Aggressiveness Detection (2021.acl-long)

Copied to clipboard

Challenge: a new method to personalize documents that are perceived differently by users is needed . a recent study found that only a few annotations of controversial documents outperform classic methods .
Approach: They propose to use some known, most controversial texts whose offensiveness is very ambiguous . they use user conformity-based measures or embeddings of their previous annotations to improve personalized reasoning .
Outcome: The proposed methods outperform standard methods in document controversy and user nonconformity . the more controversial the content, the greater the gain, the authors say .
Hate Personified: Investigating the role of LLMs in content moderation (2024.emnlp-main)

Copied to clipboard

Challenge: Our work provides preliminary guidelines and highlights the nuances of applying Large Language models in culturally sensitive cases.
Approach: They propose to use large language models to help with content moderation to assess how well the needs of diverse groups are reflected in annotated posts.
Outcome: The proposed model is able to leverage community-based flagging efforts and exposure to adversaries.
KOAS: Korean Text Offensiveness Analysis System (2021.emnlp-demo)

Copied to clipboard

Challenge: morphological richness and complex syntax of Korean cause difficulties in neural model training.
Approach: They propose a system that exploits contextual and linguistic features and estimates an offensiveness score for a Korean text.
Outcome: The proposed system exploits both contextual and linguistic features and estimates an offensiveness score for a Korean text.
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Adapting cultural values in Large Language Models presents significant challenges due to biases and data limitations.
Approach: They propose to augment World Values Survey (WVS) data with encyclopedic and scenario-based cultural narratives from Wikipedia and NormAd to address these limitations.
Outcome: The proposed approach enhances cultural distinctiveness and improves classification performance across cultures.
Can Language Models Reason about Individualistic Human Values and Preferences? (2025.acl-long)

Copied to clipboard

Challenge: Existing methods and evaluation frameworks for achieving pluralistic alignment are limited by the diversity of people, which is pre-specified and coarsely categorized, papering over individuality.
Approach: They propose to use a dataset transformed from the influential World Values Survey to study language models on the specific challenge of individualistic value reasoning.
Outcome: The proposed model can predict individualistic values with accuracies between 55% and 65%, while a precise description of individualistic value judgments cannot be approximated only via demographic information.
An Exploratory Analysis of the Relation between Offensive Language and Mental Health (2021.findings-acl)

Copied to clipboard

Challenge: Using computational models, the use of offensive language is pervasive in social media . a popular line of research is the study of machine learning classifiers to identify offensive content online .
Approach: They analyze social media posts written by individuals with depression and those without . they train computational models to compare use of offensive language with depression detection .
Outcome: The proposed models show that offensive language is more frequently used in the samples written by individuals with depression and those showing signs of depression.
Detecting Community Sensitive Norm Violations in Online Conversations (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing efforts to identify unacceptable behavior have focused on toxicity as the sole form of community norm violation.
Approach: They propose a dataset that focuses on a more complete spectrum of community norms and their violations in local conversational and global contexts.
Outcome: The proposed model improves the detection of community norm violations in local conversational and global contexts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations