The Corpus AIKIA: Using Ranking Annotation for Offensive Language Detection in Modern Greek (2024.lrec-main)
Copied to clipboard
Stella Markantonatou, Vivian Stamou, Christina Christodoulou, Georgia Apostolopoulou, Antonis Balas, George Ioannakis
| Challenge: | OLD is a less-resourced language regarding OLD. |
| Approach: | They propose to annotate OLD in Modern Greek using the lexicon of offensive terms that originates from HurtLex. |
| Outcome: | The proposed corpus is based on the lexicon of offensive terms that originates from HurtLex and can be used to detect offensive language in modern Greek. |
Similar Papers
Offensive Language Identification in Greek (2020.lrec-1)
Copied to clipboard
| Challenge: | a gap in the literature on offensive language has been addressed with studies on Spanish, Hindi, and German. |
| Approach: | They present a Greek annotated dataset for offensive language identification . it contains 4,779 tweets annotating offensive and not offensive posts from Twitter . they evaluate several computational models trained and tested on the dataset . |
| Outcome: | The proposed dataset contains 4,779 tweets annotated as offensive and not offensive . the authors show that the proposed dataset is similar to the OLID dataset for English . |
A Corpus of Turkish Offensive Language on Social Media (2020.lrec-1)
Copied to clipboard
| Challenge: | Identifying abusive, offensive, aggressive or in general inappropriate language has recently attracted interest of researchers from academic as well as commercial institutions. |
| Approach: | They propose to classify Turkish offensive language corpus using state-of-the-art annotation methods . they find 19 % of tweets contain some type of offensive language . |
| Outcome: | The proposed corpus of Turkish offensive language is the first of its kind in the world . the results show that 19 % of the tweets contain some type of offensive language . |
OLEA: Tool and Infrastructure for Offensive Language Error Analysis in English (2023.eacl-demo)
Copied to clipboard
| Challenge: | State-of-the-art models for identifying offensive language fail to generalize over nuanced or implicit cases of offensive and hateful language. |
| Approach: | They propose an open-source Python library for error analysis in the context of offensive language detection. |
| Outcome: | OLEA provides tools for error analysis in the context of detecting offensive language in English. |
Hate-Speech and Offensive Language Detection in Roman Urdu (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing research on hate-speech and offensive language detection in social media content is mainly focused on the English language. |
| Approach: | They propose to use an annotated dataset to detect hate-speech and offensive language in social media content . they propose to transfer five existing embedding models to Roman Urdu to test their performance . |
| Outcome: | The proposed model outperforms existing methods on RUHSOLD dataset and train domain-specific embeddings on more than 4.7 million tweets. |
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)
Copied to clipboard
| Challenge: | Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us . |
| Approach: | They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages. |
| Outcome: | The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning. |
A Dataset of Offensive Language in Kosovo Social Media (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media are a central part of people’s lives but are rife with bullying and offensive language, creating an unsafe environment for their users. |
| Approach: | They propose to use user-generated comments on Facebook and YouTube from selected Kosovo news platforms to annotate offensive language in Albanian. |
| Outcome: | The proposed system improves on Danish but not Albanian, on offensive language recognition and distinguishing targeted and untargeted offence. |
CoRoSeOf - An Annotated Corpus of Romanian Sexist and Offensive Tweets (2022.lrec-1)
Copied to clipboard
| Challenge: | Using CoRoSeOf, we manually annotate social media for sexist and offensive language. |
| Approach: | They introduce a large corpus of Romanian social media manually annotated for sexist and offensive language. |
| Outcome: | The proposed corpus contains 39 245 tweets annotated by multiple annotators with an agreement rate of Fleiss’= 0.45 . |
AGILe: The First Lemmatizer for Ancient Greek Inscriptions (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing models for ancient Greek inscriptions are not performant on epigraphic data due to language differences . a lemmatizer for ancient inscription data can enable meaningful generalizations, we show . |
| Approach: | They propose to train an automatic lemmatizer for ancient Greek inscriptions with 80% accuracy . they also show that existing models are not performant on epigraphic data . |
| Outcome: | The proposed model achieves above 80% accuracy on epigraphic data, and makes it available to the community. |
An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)
Copied to clipboard
| Challenge: | a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion . |
| Approach: | They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity . |
| Outcome: | The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity. |
On the Robustness of Offensive Language Classifiers (2022.acl-long)
Copied to clipboard
| Challenge: | Existing studies on offensive language classifiers have focused on primitive attacks such as misspellings and extraneous spaces. |
| Approach: | They analyze the robustness of offensive language classifiers against crafty adversarial attacks that leverage greedy- and attention-based word selection and context-aware embeddings for word replacement. |
| Outcome: | The proposed classifiers are robust against more crafty attacks that leverage greedy- and attention-based word selection and context-aware embeddings for word replacement. |