Challenge: Patronizing and Condescending Language (PCL) is a subtle but harmful type of discourse.
Approach: They propose to pre-train PCL detection models on other NLP tasks to improve their detection . they find that performance gains are possible when pre-training on sentiment, harmful language and commonsense morality.
Outcome: The proposed models improve on pre-training on other NLP tasks focusing on sentiment, harmful language and commonsense morality, compared with tasks concentrating on political speech and social justice, the authors show .

Similar Papers

PclGPT: A Large Language Model for Patronizing and Condescending Language Detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Patronizing and condescending language is an essential branch of toxic language . pre-trained language models perform poorly in detecting PCL due to its implicit toxicity traits .
Approach: They propose a novel LLM benchmark for patronizing and condescending language . they use a dataset to analyze the toxicity of patronizing condescending languages .
Outcome: The proposed model can detect patronizing and condescending language (PCL) the model can be used to analyze the toxicity of the language and to improve the detection.
Don’t Patronize Me! An Annotated Dataset with Patronizing and Condescending Language towards Vulnerable Communities (2020.coling-main)

Copied to clipboard

Challenge: a new dataset is proposed to help develop NLP models to categorize language that is patronizing or condescending towards vulnerable communities.
Approach: They propose to annotate a dataset to help develop NLP models to categorize language that is patronizing or condescending towards vulnerable communities.
Outcome: The proposed dataset supports the development of NLP models to categorize language that is patronizing or condescending towards vulnerable communities.
Probing Toxic Content in Large Pre-Trained Language Models (2021.acl-long)

Copied to clipboard

Challenge: Existing studies on pre-trained language models have shown that they carry harmful biases towards different social groups.
Approach: They propose a method to probe English, French, and Arabic PTLMs and quantify the potentially harmful content they convey with respect to a set of templates.
Outcome: The proposed method analyzes PTLMs to predict masked tokens at the end of sentences to assess their toxicity.
Do Neural Language Models Overcome Reporting Bias? (2020.coling-main)

Copied to clipboard

Challenge: Recent studies show that pre-trained language models can overcome reporting bias by estimating the plausibility of rare but unspoken facts.
Approach: They revisit the experiments conducted by Gordon and Van Durme (2013) . they find that pre-trained language models overestimate the very rare .
Outcome: The proposed approach overestimates the rare at the expense of the rare, while minimizing reporting bias.
Revisiting Implicitly Abusive Language Detection: Evaluating LLMs in Zero-Shot and Few-Shot Settings (2025.coling-main)

Copied to clipboard

Challenge: Current research focuses on explicit abusive language, but subtler forms of IAL remain insufficiently studied.
Approach: They evaluate the models' capabilities in classifying sentences directly as either IAL or benign, and in extracting linguistic features associated with IAL.
Outcome: The proposed models outperform the best previously reported methods in classifying sentences directly as IAL or benign and extracting linguistic features associated with IAL.
Implicitly Abusive Language – What does it actually look like and why are we not getting there? (2021.naacl-main)

Copied to clipboard

Challenge: Existing datasets make learning implicit abuse difficult, argues a new position paper . a lack of work on implicit abuse has limited the effectiveness of automatic detection .
Approach: They argue that existing datasets make learning implicit abuse difficult . they propose a divide-and-conquer strategy to detect implicit abuse .
Outcome: The proposed model could be improved to detect implicit abuse in a dataset with a standardized model.
Euphemistic Abuse – A New Dataset and Classification Experiments for Implicitly Abusive Language (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web.
Approach: They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances.
Outcome: The proposed classifier augments training data with automatically-generated GPT-3 completions.
TalkDown: A Corpus for Condescension Detection in Context (D19-1)

Copied to clipboard

Challenge: condescending language use can bring dialogues to an end and disrupt healthy communities.
Approach: They propose a model that uses a language-only model to model condescending linguistic acts in context.
Outcome: a new model of condescending language use improves performance and motivates techniques . the model can estimate condescension rates in various online communities and relate these differences to community norms .
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models (2023.acl-long)

Copied to clipboard

Challenge: Hundreds of studies have highlighted ethical issues in NLP models .
Approach: They propose to measure media biases in LMs trained on diverse data sources . they focus on hate speech and misinformation detection .
Outcome: The proposed methods quantify the fairness of downstream NLP models trained on politically biased LMs.
ConPrompt: Pre-training a Language Model with Machine-Generated Data for Implicit Hate Speech Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing pre-trained language models for hate speech detection are not specialized in implicit hate speech.
Approach: They propose a pre-trained language model for implicit hate speech detection that leverages machine-generated data to train the model.
Outcome: The proposed model can be trained on a massive hate speech dataset with positive samples . it can be generalized and reduce identity term bias, the authors show .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations