From legal to technical concept: Towards an automated classification of German political Twitter postings as criminal offenses (N19-1)
Copied to clipboard
| Challenge: | 'Network Enforcement Act' provides for a regulatory framework for 'illegal content' on social network platforms like Twitter or Facebook. |
| Approach: | They propose a data annotation schema to determine whether a particular tweet could constitute a criminal offense and a binary classification schema to help with this. |
| Outcome: | The proposed schema shows that the majority of offensive posts do not constitute a criminal offense and still contribute to public discourse. |
Similar Papers
A Dataset of Offensive German Language Tweets Annotated for Speech Acts (2022.lrec-1)
Copied to clipboard
| Challenge: | Using speech act analysis, we analysed 600 offensive and non-offensive tweets in germany . a large body of research exists on the pragmatic characteristics of offensive language . |
| Approach: | They analyze German offensive and non-offensive tweets and use a subset of the 2019 GermEval Shared Task on the Identification of Offensive Language dataset. |
| Outcome: | The proposed dataset includes 600 offensive and non-offensive tweets annotated for speech acts in germany. |
Predicting the Type and Target of Offensive Posts in Social Media (N19-1)
Copied to clipboard
| Challenge: | Prior work focused on detecting specific types of offensive content, such as hate speech, cyberbullying, or cyber-aggression. |
| Approach: | They propose to use a dataset to identify offensive content in social media . they compare the performance of different machine learning models to OLID . |
| Outcome: | The proposed dataset contains tweets annotated for offensive content using a fine-grained three-layer annotation scheme. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
Offensive Language and Hate Speech Detection for Danish (2020.lrec-1)
Copied to clipboard
| Challenge: | a growing number of social media platforms are detecting and dealing with offensive language . a recent study found that the best performing system for English is best for Danish . |
| Approach: | They propose automatic methods to detect offensive language on social media platforms . they use user-generated comments from various social media sites to find offensive language . |
| Outcome: | The proposed system performs best for both English and Danish language . it achieves a macro averaged F1-score of 0.74 and a best for Danish achieves 0.73 . |
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)
Copied to clipboard
| Challenge: | Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us . |
| Approach: | They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages. |
| Outcome: | The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning. |
DeFaktS: A German Dataset for Fine-Grained Disinformation Detection through Social Media Framing (2024.lrec-main)
Copied to clipboard
| Challenge: | Distinctively curated across various news topics, DeFaktS offers an unparalleled insight into disinformation’s diverse characteristics. |
| Approach: | They propose to annotate every structural component and semantic element of a news piece, eliminating the need for external knowledge sources. |
| Outcome: | The proposed dataset contains 105,855 posts with 20,008 meticulously labeled tweets and eliminates the need for external knowledge sources. |
A Corpus of Turkish Offensive Language on Social Media (2020.lrec-1)
Copied to clipboard
| Challenge: | Identifying abusive, offensive, aggressive or in general inappropriate language has recently attracted interest of researchers from academic as well as commercial institutions. |
| Approach: | They propose to classify Turkish offensive language corpus using state-of-the-art annotation methods . they find 19 % of tweets contain some type of offensive language . |
| Outcome: | The proposed corpus of Turkish offensive language is the first of its kind in the world . the results show that 19 % of the tweets contain some type of offensive language . |
An Annotated Corpus for Sexism Detection in French Tweets (2020.lrec-1)
Copied to clipboard
Patricia Chiril, Véronique Moriceau, Farah Benamara, Alda Mari, Gloria Origgi, Marlène Coulomb-Gully
| Challenge: | Social media networks allow users to share opinions and sentiments, which can cause a large spreading of hatred or abusive messages. |
| Approach: | They propose to annotate 12,000 tweets with a sexism detection scheme in France . they propose to use deep learning to detect if a message with sexist content is really s. |
| Outcome: | The proposed scheme detects sexist content and identifies if it is really sexism . the proposed scheme is the first of its kind in the u.s. |
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation. |
| Approach: | They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic . |
| Outcome: | The proposed models do not generalize, indicating heterogeneous political users. |
He said “who’s gonna take care of your children when you are at ACL?”: Reported Sexist Acts are Not Sexist (2020.acl-main)
Copied to clipboard
Patricia Chiril, Véronique Moriceau, Farah Benamara, Alda Mari, Gloria Origgi, Marlène Coulomb-Gully
| Challenge: | Sexism is prejudice or discrimination based on a person's gender. |
| Approach: | They propose to use a French dataset annotated for sexism detection to characterize sexist content and to train deep learning experiments on tweets. |
| Outcome: | The proposed dataset is the first to be used for sexism detection in France and constitutes a first step towards offensive content moderation. |