Challenge: Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals.
Approach: They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments .
Outcome: The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube.

Similar Papers

So Hateful! Building a Multi-Label Hate Speech Annotated Arabic Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Social media enables widespread propagation of hate speech targeting groups based on ethnicity, religion, or other characteristics.
Approach: They analyze 70,000 Arabic tweets to identify hate speech patterns and train models . 15% of tweets contain offensive language while 6% have hate speech . authors hope to prevent spread of hateful content on social media platforms .
Outcome: The analysis of 70,000 Arabic tweets shows that 15% of tweets contain offensive language while 6% have hate speech . 10% of tweet provide verifiable factual claims, and 7% are deemed important .
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)

Copied to clipboard

Challenge: Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us .
Approach: They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages.
Outcome: The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning.
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)

Copied to clipboard

Challenge: 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms.
Approach: They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
Outcome: The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English.
Toxic Language Detection in Social Media for Brazilian Portuguese: New Dataset and Multilingual Analysis (2020.aacl-main)

Copied to clipboard

Challenge: Hate speech and toxic comments are a common concern of social media platform users . identifying toxic comments is important for studying and preventing the proliferation of toxicity in social media.
Approach: They propose to use Brazilian Portuguese to analyze toxic or non-toxic tweets . they propose to analyze tweets as toxic or in different types of toxicity .
Outcome: The proposed model achieves 76% macro-F1 score using monolingual data in the binary case.
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes.
Approach: They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors.
Outcome: The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups.
MARASTA: A Multi-dialectal Arabic Cross-domain Stance Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Approximately half of the sentences are in Modern Standard Arabic (MSA) for each region, and the other half is in the region’s respective dialect.
Approach: They propose a cross-domain and multi-dialectal stance corpus for Arabic that includes four regions in the Arab World and covers the main Arabic dialect groups.
Outcome: The proposed corpus outperforms the state-of-the-art dataset in stance detection and dialect and dialect classes.
Hate-Speech and Offensive Language Detection in Roman Urdu (2020.emnlp-main)

Copied to clipboard

Challenge: Existing research on hate-speech and offensive language detection in social media content is mainly focused on the English language.
Approach: They propose to use an annotated dataset to detect hate-speech and offensive language in social media content . they propose to transfer five existing embedding models to Roman Urdu to test their performance .
Outcome: The proposed model outperforms existing methods on RUHSOLD dataset and train domain-specific embeddings on more than 4.7 million tweets.
CONAN - COunter NArratives through Nichesourcing: a Multilingual Dataset of Responses to Fight Online Hate Speech (P19-1)

Copied to clipboard

Challenge: Davidson et al., 2017): social media platforms and governmental organizations have taken steps to tackle hate speech . Davidson and Norton, 2017: a dataset of hate-speech/counter-narrative pairs is created . authors: identifying hate speech is challenging for the broadness and nuances in cultures and languages .
Approach: They propose to build a large-scale, multilingual, expert-based dataset of hate-speech/counter-narrative pairs . they provide additional annotations about expert demographics, hate and response type .
Outcome: The proposed dataset provides an analysis of hate-speech/counter-narrative pairs in three languages.
Multilingual Offensive Language Identification with Cross-lingual Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Several studies investigating methods to detect offensive content in social media use English data.
Approach: They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources.
Outcome: The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish.
DART: A Large Dataset of Dialectal Arabic Tweets (L18-1)

Copied to clipboard

Challenge: The Arabic language is the fifth most widely spoken language in the world; more than 380 million people speak and write in Arabic.
Approach: They propose to build a large manually-annotated multi-dialect dataset of Arabic tweets that is publicly available.
Outcome: The proposed dataset is well-balanced over five main Arabic dialects: Egyptian, Maghrebi, Levantine, Gulf, and Iraqi.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations