Challenge: Existing tools for quantifying incivility online, in news and in congressional debates are inadequate for the analysis of incivility in news.
Approach: They develop a Jigsaw Perspective API to quantify incivility in news . they show that toxicity models are inadequate for the analysis of incivility in news.
Outcome: The Jigsaw Perspective API detects incivility on a corpus of American news articles.

Similar Papers

A Closer Look at Multidimensional Online Political Incivility (2024.emnlp-main)

Copied to clipboard

Challenge: 80% of the uncivil tweets are authored by 20% of the users, where users who are politically engaged are more inclined to use uncival language.
Approach: They analysed 13K political tweets in the U.S. using crowd sourcing and classified them by their respective categories.
Outcome: The proposed method enables us to characterise the distribution of incivility across users and geopolitical regions.
The Computational Anatomy of Humility: Modeling Intellectual Humility in Online Public Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: enhancing the quality of online public discourse requires promoting foundational human virtues, such as “intellectual humility” (IH) . discourse on social media rewards forgetting our virtuous selves, embedding users within echo chambers and causing negative affect towards those who hold different beliefs.
Approach: They propose to use a codebook to measure "intellectual humility" they manually validated the codebook and used it to develop LLM-based models .
Outcome: The proposed model achieves a Macro-F1 score of 0.64 across labels and 0.70 when predicting IH/IA/Neutral at the coarse level.
“It’s Not Just Hate”: A Multi-Dimensional Perspective on Detecting Harmful Speech Online (2022.emnlp-main)

Copied to clipboard

Challenge: Detecting offensive content is becoming a critical task in natural language processing . but most datasets use a single binary label for hate or incivility, even though each concept is multi-faceted . a more fine-grained multi-label approach addresses conceptual and performance issues .
Approach: They propose to use a dataset to annotate offensive online speech with six labels . they propose to apply a more fine-grained approach to predicting incivility and hateful content .
Outcome: The proposed approach outperforms or matches benchmark datasets on the annotated tweets.
Finding Microaggressions in the Wild: A Case for Locating Elusive Phenomena in Social Media Posts (D19-1)

Copied to clipboard

Challenge: Existing tools for hate speech detection and sentiment analysis cannot detect veiled offensiveness of microaggressions . linguistic subtlety of micro-aggressives has made it difficult to analyze their exact nature .
Approach: They propose a typology of microaggressions based on a subset of data . they propose an objective criterion for annotation and an active-learning procedure .
Outcome: The proposed typology of microaggressions is based on a subset of social media data.
Discovering Biased News Articles Leveraging Multiple Human Annotations (2020.lrec-1)

Copied to clipboard

Challenge: Political propaganda and one-sided views can be found in the news and can cause distrust in media.
Approach: They propose to annotate politically biased news articles by an algorithm annotated by domain experts and crowd workers and to compare them to crowd workers.
Outcome: The proposed method compares domain experts to crowd workers and shows that bias can be detected automatically.
On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research (2023.emnlp-main)

Copied to clipboard

Challenge: Perception of toxicity evolves over time and differs between geographies and cultural backgrounds.
Approach: They propose to use a more structured approach to evaluating toxicity over time . they suggest that research that relied on automatic toxicity scores may have resulted in inaccurate results.
Outcome: The Perspective API has been updated to reflect the changes in toxicity scores.
Toxicity Detection: Does Context Really Matter? (2020.acl-main)

Copied to clipboard

Challenge: Existing ‘toxicity’ detection datasets and models ignore the context of the posts, implicitly assuming that comments may be judged independently.
Approach: They limit the notion of context to the previous post in the thread and the discussion title and focus on how it affects human judgement.
Outcome: The proposed model can amplify or mitigate perceived toxicity of posts and a small but significant subset of manually labeled posts end up having the opposite toxicity labels if the annotators are not provided with context.
Toxicity, Morality, and Speech Act Guided Stance Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies that focus on stance detection ignore the speech act, toxic, and moral features of tweets or lack an efficient architecture to detect the attitudes across targets.
Approach: They propose a multitasking model that extracts valence, arousal, and dominance aspects hidden in tweets and injects the emotional sense into the embedded text followed by an efficient attention framework to correctly detect the tweet’s stance.
Outcome: The proposed model exploits the toxicity, morality, and speech act features of the tweets to detect the public's stance.
A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis (2026.findings-acl)

Copied to clipboard

Challenge: a large-scale label set for media outlets from Media Bias/Fact Check (MBFC) is lacking in the field.
Approach: They propose to use a large-scale label set to analyze outlets' representations . they also propose to evaluate embedding views and fusion strategies .
Outcome: The proposed method achieves state-of-the-art results on ACL-2020 and establishes strong benchmarks on MBFC-2025.
Detecting Community Sensitive Norm Violations in Online Conversations (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing efforts to identify unacceptable behavior have focused on toxicity as the sole form of community norm violation.
Approach: They propose a dataset that focuses on a more complete spectrum of community norms and their violations in local conversational and global contexts.
Outcome: The proposed model improves the detection of community norm violations in local conversational and global contexts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations