A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection (2020.lrec-1)
Copied to clipboard
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-gyo Jung, Bernard J. Jansen, Joni Salminen
| Challenge: | Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals. |
| Approach: | They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments . |
| Outcome: | The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube. |
Similar Papers
So Hateful! Building a Multi-Label Hate Speech Annotated Arabic Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Social media enables widespread propagation of hate speech targeting groups based on ethnicity, religion, or other characteristics. |
| Approach: | They analyze 70,000 Arabic tweets to identify hate speech patterns and train models . 15% of tweets contain offensive language while 6% have hate speech . authors hope to prevent spread of hateful content on social media platforms . |
| Outcome: | The analysis of 70,000 Arabic tweets shows that 15% of tweets contain offensive language while 6% have hate speech . 10% of tweet provide verifiable factual claims, and 7% are deemed important . |
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)
Copied to clipboard
| Challenge: | Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us . |
| Approach: | They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages. |
| Outcome: | The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning. |
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)
Copied to clipboard
Ritesh Kumar, Shyam Ratan, Siddharth Singh, Enakshi Nandi, Laishram Niranjana Devi, Akash Bhagat, Yogesh Dawer, Bornini Lahiri, Akanksha Bansal, Atul Kr. Ojha
| Challenge: | 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms. |
| Approach: | They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. |
| Outcome: | The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English. |
Toxic Language Detection in Social Media for Brazilian Portuguese: New Dataset and Multilingual Analysis (2020.aacl-main)
Copied to clipboard
| Challenge: | Hate speech and toxic comments are a common concern of social media platform users . identifying toxic comments is important for studying and preventing the proliferation of toxicity in social media. |
| Approach: | They propose to use Brazilian Portuguese to analyze toxic or non-toxic tweets . they propose to analyze tweets as toxic or in different types of toxicity . |
| Outcome: | The proposed model achieves 76% macro-F1 score using monolingual data in the binary case. |
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. |
| Approach: | They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors. |
| Outcome: | The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups. |
MARASTA: A Multi-dialectal Arabic Cross-domain Stance Corpus (2024.lrec-main)
Copied to clipboard
Anis Charfi, Mabrouka Ben-Sghaier, Andria Samy Raouf Atalla, Raghda Akasheh, Sara Al-Emadi, Wajdi Zaghouani
| Challenge: | Approximately half of the sentences are in Modern Standard Arabic (MSA) for each region, and the other half is in the region’s respective dialect. |
| Approach: | They propose a cross-domain and multi-dialectal stance corpus for Arabic that includes four regions in the Arab World and covers the main Arabic dialect groups. |
| Outcome: | The proposed corpus outperforms the state-of-the-art dataset in stance detection and dialect and dialect classes. |
Hate-Speech and Offensive Language Detection in Roman Urdu (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing research on hate-speech and offensive language detection in social media content is mainly focused on the English language. |
| Approach: | They propose to use an annotated dataset to detect hate-speech and offensive language in social media content . they propose to transfer five existing embedding models to Roman Urdu to test their performance . |
| Outcome: | The proposed model outperforms existing methods on RUHSOLD dataset and train domain-specific embeddings on more than 4.7 million tweets. |
CONAN - COunter NArratives through Nichesourcing: a Multilingual Dataset of Responses to Fight Online Hate Speech (P19-1)
Copied to clipboard
| Challenge: | Davidson et al., 2017): social media platforms and governmental organizations have taken steps to tackle hate speech . Davidson and Norton, 2017: a dataset of hate-speech/counter-narrative pairs is created . authors: identifying hate speech is challenging for the broadness and nuances in cultures and languages . |
| Approach: | They propose to build a large-scale, multilingual, expert-based dataset of hate-speech/counter-narrative pairs . they provide additional annotations about expert demographics, hate and response type . |
| Outcome: | The proposed dataset provides an analysis of hate-speech/counter-narrative pairs in three languages. |
Multilingual Offensive Language Identification with Cross-lingual Embeddings (2020.emnlp-main)
Copied to clipboard
| Challenge: | Several studies investigating methods to detect offensive content in social media use English data. |
| Approach: | They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources. |
| Outcome: | The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish. |
DART: A Large Dataset of Dialectal Arabic Tweets (L18-1)
Copied to clipboard
| Challenge: | The Arabic language is the fifth most widely spoken language in the world; more than 380 million people speak and write in Arabic. |
| Approach: | They propose to build a large manually-annotated multi-dialect dataset of Arabic tweets that is publicly available. |
| Outcome: | The proposed dataset is well-balanced over five main Arabic dialects: Egyptian, Maghrebi, Levantine, Gulf, and Iraqi. |