Challenge: Xenophobia and polarization have accompanied widespread social media usage in many nations, attracting many researchers.
Approach: They apply natural language processing techniques to characterize Twitter users who began to post anti-Asian hate messages during COVID-19.
Outcome: The results show that it is possible to predict who later posted anti-Asian slurs on Twitter and Reddit.

Similar Papers

A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech (2024.acl-long)

Copied to clipboard

Challenge: Using data from 420k Twitter posts, we characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection.
Approach: They develop a codebook to characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection.
Outcome: The proposed codebook analyzes 420k tweets over 3 years and compares classifiers with hateful speech classifier classifier to detect hateful content.
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter (2025.acl-long)

Copied to clipboard

Challenge: Prior work on automated hate speech detection models has been limited due to systematic biases in evaluation datasets and poor performance across geographies.
Approach: They propose to construct a global hate speech dataset representative of social media settings from tweets posted on September 21, 2022.
Outcome: The proposed dataset covers eight languages and four English-speaking countries and covers eight countries where English is the main language on Twitter.
Generating Counter Narratives against Online Hate Speech: Data and Strategies (2020.acl-main)

Copied to clipboard

Challenge: Hate Speech (HS) is a pervasive issue that spreads quickly and widely . research has focused on avoiding undesired effects that come with content moderation .
Approach: They propose to use large scale unsupervised language models to generate responses to hate effectively using large scale models.
Outcome: The proposed methods lack quality data and produce generic/repetitive responses.
Leveraging Intra-User and Inter-User Representation Learning for Automated Hate Speech Detection (N18-2)

Copied to clipboard

Challenge: Existing methods that focus on a single tweet as input are likely to yield high false positive and negative rates.
Approach: They propose a model that leverages intra-user and inter-user representation learning to improve hate speech detection on Twitter by suppressing the noise in a single Tweet.
Outcome: The proposed model significantly improves the f-score of a strong bidirectional LSTM model by 10.1%.
COVID-19 and Misinformation: A Large-Scale Lexical Analysis on Twitter (2021.acl-srw)

Copied to clipboard

Challenge: Social media is used by individuals and organisations as a platform to spread misinformation.
Approach: They compile a large corpus of tweets related to coronavirus and perform an analysis to discover patterns with respect to vocabulary usage.
Outcome: The proposed model based on lexical features is effective in identifying misinformation-related tweets with accuracy over 80%.
Predicting the Type and Target of Offensive Posts in Social Media (N19-1)

Copied to clipboard

Challenge: Prior work focused on detecting specific types of offensive content, such as hate speech, cyberbullying, or cyber-aggression.
Approach: They propose to use a dataset to identify offensive content in social media . they compare the performance of different machine learning models to OLID .
Outcome: The proposed dataset contains tweets annotated for offensive content using a fine-grained three-layer annotation scheme.
Stance Detection in COVID-19 Tweets (2021.acl-long)

Copied to clipboard

Challenge: a global pandemic of COVID-19 has forced major changes in our daily lives . a new stance detection dataset is being used to track the stances of Twitter users .
Approach: They use Twitter stance data to collect stances on topics related to the pandemic . they train models to take advantage of large amounts of unlabeled data .
Outcome: The proposed model improves on existing stance detection datasets and unlabeled data.
A Dataset for Investigating the Impact of Context for Offensive Language Detection in Tweets (2023.findings-emnlp)

Copied to clipboard

Challenge: Offensive language detection is crucial in natural language processing . we investigated the importance of contextual information for detecting offensive language in tweets .
Approach: They investigated the importance of contextual information for detecting offensive language in tweets . they used a Turkish tweet dataset with over 28,000 tweet-reply pairs .
Outcome: The proposed model performs better with and without contextual information than with and with contextual information.
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)

Copied to clipboard

Challenge: a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms.
Approach: They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia .
Outcome: The proposed dataset shows that it is useful in monolingual vs. multilingual settings.
Hate Speech and Offensive Language Detection in Bengali (2022.aacl-main)

Copied to clipboard

Challenge: Existing research on hate speech detection in English does not cover low-resource languages like Bengali.
Approach: They develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets.
Outcome: The proposed model outperforms other models on training actual and romanized datasets by interpreting the semantic expressions better.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations