Polarized Opinion Detection Improves the Detection of Toxic Language (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing methods for estimating polarized annotations are un-normalized and difficult to exploit in machine learning. |
| Approach: | They propose a method for K-class text classification that exploits polarized texts in the dataset. |
| Outcome: | The proposed method exploits polarized texts in a dataset and can improve classification performance. |
Similar Papers
Contrastive Learning as a Polarizer: Mitigating Gender Bias by Fair and Biased sentences (2024.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have highlighted social biases inherent in training data can lead models to learn and propagate them. |
| Approach: | They propose a contrastive learning method that uses anchor points to push further negatives and pull closer positives within the representation space. |
| Outcome: | The proposed method achieves state-of-the-art in the ICAT score on the StereoSet, a benchmark for measuring bias in models. |
A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information (2025.findings-acl)
Copied to clipboard
Lucky Susanto, Musa Izzanardi Wijanarko, Prasetia Anugrah Pratama, Zilu Tang, Fariz Akyas, Traci Hong, Ika Karlina Idris, Alham Fikri Aji, Derry Tanti Wijaya
| Challenge: | Prior research has focused on toxicity and polarization as separate problems . extreme polarizing deepens divisions, often leading to hostility and fragmentation . |
| Approach: | They propose to use a multi-label Indonesian dataset annotated for toxicity, polarization, and annotator demographic information to study polarizing language and toxicity. |
| Outcome: | The proposed dataset shows that polarization cues improve toxicity classification and vice versa. |
Challenges in Automated Debiasing for Toxic Language Detection (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for debiasing toxic language data are limited in their ability to prevent biased behavior in toxic language detection systems. |
| Approach: | They propose to debiase toxic language detection models using lexical and dialectal markers using synthetic labels instead of traditional methods. |
| Outcome: | The proposed method reduces dialectal associations with toxicity despite the use of synthetic labels . |
Why Do Document-Level Polarity Classifiers Fail? (2021.naacl-main)
Copied to clipboard
| Challenge: | a new method to characterize, quantify and measure the impact of hard instances is proposed . a method to label hard instances can shed light on why and when classifiers fail, authors say . |
| Approach: | They propose a method to characterize, quantify and measure the impact of hard instances in polarity classification of movie reviews. |
| Outcome: | The proposed method can quantify the impact of hard instances in polarity classification . it can shed light on why and when classifiers fail, the authors say . |
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection (2025.findings-naacl)
Copied to clipboard
Tomáš Horych, Christoph Mandl, Terry Ruas, Andre Greiner-Petter, Bela Gipp, Akiko Aizawa, Timo Spinde
| Challenge: | Recent research suggests using Large Language Models (LLMs) to automate the annotation process, reducing these costs while maintaining data quality. |
| Approach: | They propose to use Large Language Models to automate annotation process and train classifiers on large datasets. |
| Outcome: | The proposed model outperforms all of the annotator LLMs on two media bias benchmark datasets (BABE and BASIL) while maintaining data quality. |
Inference Annotation of a Chinese Corpus for Opinion Mining (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing tools for opinion mining can accurately predict the writer's attitude in simple explicit sentences. |
| Approach: | They propose to define inference, classify different types and provide an annotation framework to analyze the annotation results. |
| Outcome: | The proposed framework defines inference type, polarity and topic and analyzes the results. |
Are Text Classifiers Xenophobic? A Country-Oriented Bias Detection Method with Least Confounding Variables (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for detecting biases are biased because of confounding variables . authors propose a method to detect the biased classifier on any type of unlabeled data . |
| Approach: | They propose a method to detect biases of a specific fine-tuned classifier on unlabeled data. |
| Outcome: | The proposed method detects biases on unlabeled data on named entity perturbations . it uses name-entity recognition on target-domain data and morphosynctactically different languages spoken in relation to countries of the target groups . |
KoSBI: A Dataset for Mitigating Social Bias Risks Towards Safer Large Language Model Applications (2023.acl-industry)
Copied to clipboard
| Challenge: | Existing research and resources are not readily applicable in South Korea due to the differences in language and culture, both of which significantly affect the biases and targeted demographic groups. |
| Approach: | They propose a social bias dataset of 34k pairs of contexts and sentences in Korean covering 72 demographic groups in 15 categories. |
| Outcome: | The proposed dataset reduces social biases by 16.47%p on average for HyperClova (30B and 82B), and GPT-3. |
“Why do I feel offended?” - Korean Dataset for Offensive Language Identification (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods for detecting offensive content rely on labeled datasets, but few consider low-resource languages with relatively less data available for training. |
| Approach: | They propose to use Korean as a dataset for offensive language identification . they propose to perform abusive language detection and sentiment analysis to help identify offensive languages. |
| Outcome: | The proposed datasets improve the performance of offensive language identification in Korean, while the existing methods are limited. |
A Weakly Supervised Classifier and Dataset of White Supremacist Language (2023.acl-short)
Copied to clipboard
| Challenge: | Existing studies on white supremacist language have focused on specific hateful ideologies, but little attention has been given to specific hate speech. |
| Approach: | They propose a weakly supervised classifier for detecting white supremacist language . they use large datasets of white supremacy domains paired with neutral and anti-racist data from similar domains to train the classifiers. |
| Outcome: | The proposed classifiers outperform previous studies on white supremacist classification on unseen datasets and find strong generalization performance for models with weakly annotated data. |