Challenge: 'Network Enforcement Act' provides for a regulatory framework for 'illegal content' on social network platforms like Twitter or Facebook.
Approach: They propose a data annotation schema to determine whether a particular tweet could constitute a criminal offense and a binary classification schema to help with this.
Outcome: The proposed schema shows that the majority of offensive posts do not constitute a criminal offense and still contribute to public discourse.

Similar Papers

A Dataset of Offensive German Language Tweets Annotated for Speech Acts (2022.lrec-1)

Copied to clipboard

Challenge: Using speech act analysis, we analysed 600 offensive and non-offensive tweets in germany . a large body of research exists on the pragmatic characteristics of offensive language .
Approach: They analyze German offensive and non-offensive tweets and use a subset of the 2019 GermEval Shared Task on the Identification of Offensive Language dataset.
Outcome: The proposed dataset includes 600 offensive and non-offensive tweets annotated for speech acts in germany.
Predicting the Type and Target of Offensive Posts in Social Media (N19-1)

Copied to clipboard

Challenge: Prior work focused on detecting specific types of offensive content, such as hate speech, cyberbullying, or cyber-aggression.
Approach: They propose to use a dataset to identify offensive content in social media . they compare the performance of different machine learning models to OLID .
Outcome: The proposed dataset contains tweets annotated for offensive content using a fine-grained three-layer annotation scheme.
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
Offensive Language and Hate Speech Detection for Danish (2020.lrec-1)

Copied to clipboard

Challenge: a growing number of social media platforms are detecting and dealing with offensive language . a recent study found that the best performing system for English is best for Danish .
Approach: They propose automatic methods to detect offensive language on social media platforms . they use user-generated comments from various social media sites to find offensive language .
Outcome: The proposed system performs best for both English and Danish language . it achieves a macro averaged F1-score of 0.74 and a best for Danish achieves 0.73 .
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)

Copied to clipboard

Challenge: Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us .
Approach: They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages.
Outcome: The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning.
DeFaktS: A German Dataset for Fine-Grained Disinformation Detection through Social Media Framing (2024.lrec-main)

Copied to clipboard

Challenge: Distinctively curated across various news topics, DeFaktS offers an unparalleled insight into disinformation’s diverse characteristics.
Approach: They propose to annotate every structural component and semantic element of a news piece, eliminating the need for external knowledge sources.
Outcome: The proposed dataset contains 105,855 posts with 20,008 meticulously labeled tweets and eliminates the need for external knowledge sources.
A Corpus of Turkish Offensive Language on Social Media (2020.lrec-1)

Copied to clipboard

Challenge: Identifying abusive, offensive, aggressive or in general inappropriate language has recently attracted interest of researchers from academic as well as commercial institutions.
Approach: They propose to classify Turkish offensive language corpus using state-of-the-art annotation methods . they find 19 % of tweets contain some type of offensive language .
Outcome: The proposed corpus of Turkish offensive language is the first of its kind in the world . the results show that 19 % of the tweets contain some type of offensive language .
An Annotated Corpus for Sexism Detection in French Tweets (2020.lrec-1)

Copied to clipboard

Challenge: Social media networks allow users to share opinions and sentiments, which can cause a large spreading of hatred or abusive messages.
Approach: They propose to annotate 12,000 tweets with a sexism detection scheme in France . they propose to use deep learning to detect if a message with sexist content is really s.
Outcome: The proposed scheme detects sexist content and identifies if it is really sexism . the proposed scheme is the first of its kind in the u.s.
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation.
Approach: They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic .
Outcome: The proposed models do not generalize, indicating heterogeneous political users.
He said “who’s gonna take care of your children when you are at ACL?”: Reported Sexist Acts are Not Sexist (2020.acl-main)

Copied to clipboard

Challenge: Sexism is prejudice or discrimination based on a person's gender.
Approach: They propose to use a French dataset annotated for sexism detection to characterize sexist content and to train deep learning experiments on tweets.
Outcome: The proposed dataset is the first to be used for sexism detection in France and constitutes a first step towards offensive content moderation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations