Predicting Anti-Asian Hateful Users on Twitter during COVID-19 (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Xenophobia and polarization have accompanied widespread social media usage in many nations, attracting many researchers. |
| Approach: | They apply natural language processing techniques to characterize Twitter users who began to post anti-Asian hate messages during COVID-19. |
| Outcome: | The results show that it is possible to predict who later posted anti-Asian slurs on Twitter and Reddit. |
Similar Papers
A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech (2024.acl-long)
Copied to clipboard
Gaurav Verma, Rynaa Grover, Jiawei Zhou, Binny Mathew, Jordan Kraemer, Munmun Choudhury, Srijan Kumar
| Challenge: | Using data from 420k Twitter posts, we characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection. |
| Approach: | They develop a codebook to characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection. |
| Outcome: | The proposed codebook analyzes 420k tweets over 3 years and compares classifiers with hateful speech classifier classifier to detect hateful content. |
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter (2025.acl-long)
Copied to clipboard
Manuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale, Samuel Fraiberger, Victor Orozco-Olvera, Paul Röttger
| Challenge: | Prior work on automated hate speech detection models has been limited due to systematic biases in evaluation datasets and poor performance across geographies. |
| Approach: | They propose to construct a global hate speech dataset representative of social media settings from tweets posted on September 21, 2022. |
| Outcome: | The proposed dataset covers eight languages and four English-speaking countries and covers eight countries where English is the main language on Twitter. |
Generating Counter Narratives against Online Hate Speech: Data and Strategies (2020.acl-main)
Copied to clipboard
| Challenge: | Hate Speech (HS) is a pervasive issue that spreads quickly and widely . research has focused on avoiding undesired effects that come with content moderation . |
| Approach: | They propose to use large scale unsupervised language models to generate responses to hate effectively using large scale models. |
| Outcome: | The proposed methods lack quality data and produce generic/repetitive responses. |
Leveraging Intra-User and Inter-User Representation Learning for Automated Hate Speech Detection (N18-2)
Copied to clipboard
| Challenge: | Existing methods that focus on a single tweet as input are likely to yield high false positive and negative rates. |
| Approach: | They propose a model that leverages intra-user and inter-user representation learning to improve hate speech detection on Twitter by suppressing the noise in a single Tweet. |
| Outcome: | The proposed model significantly improves the f-score of a strong bidirectional LSTM model by 10.1%. |
COVID-19 and Misinformation: A Large-Scale Lexical Analysis on Twitter (2021.acl-srw)
Copied to clipboard
| Challenge: | Social media is used by individuals and organisations as a platform to spread misinformation. |
| Approach: | They compile a large corpus of tweets related to coronavirus and perform an analysis to discover patterns with respect to vocabulary usage. |
| Outcome: | The proposed model based on lexical features is effective in identifying misinformation-related tweets with accuracy over 80%. |
Predicting the Type and Target of Offensive Posts in Social Media (N19-1)
Copied to clipboard
| Challenge: | Prior work focused on detecting specific types of offensive content, such as hate speech, cyberbullying, or cyber-aggression. |
| Approach: | They propose to use a dataset to identify offensive content in social media . they compare the performance of different machine learning models to OLID . |
| Outcome: | The proposed dataset contains tweets annotated for offensive content using a fine-grained three-layer annotation scheme. |
Stance Detection in COVID-19 Tweets (2021.acl-long)
Copied to clipboard
| Challenge: | a global pandemic of COVID-19 has forced major changes in our daily lives . a new stance detection dataset is being used to track the stances of Twitter users . |
| Approach: | They use Twitter stance data to collect stances on topics related to the pandemic . they train models to take advantage of large amounts of unlabeled data . |
| Outcome: | The proposed model improves on existing stance detection datasets and unlabeled data. |
A Dataset for Investigating the Impact of Context for Offensive Language Detection in Tweets (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Offensive language detection is crucial in natural language processing . we investigated the importance of contextual information for detecting offensive language in tweets . |
| Approach: | They investigated the importance of contextual information for detecting offensive language in tweets . they used a Turkish tweet dataset with over 28,000 tweet-reply pairs . |
| Outcome: | The proposed model performs better with and without contextual information than with and with contextual information. |
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)
Copied to clipboard
Firoj Alam, Shaden Shaar, Fahim Dalvi, Hassan Sajjad, Alex Nikolov, Hamdy Mubarak, Giovanni Da San Martino, Ahmed Abdelali, Nadir Durrani, Kareem Darwish, Abdulaziz Al-Homaid, Wajdi Zaghouani, Tommaso Caselli, Gijs Danoe, Friso Stolk, Britt Bruntink, Preslav Nakov
| Challenge: | a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms. |
| Approach: | They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia . |
| Outcome: | The proposed dataset shows that it is useful in monolingual vs. multilingual settings. |
Hate Speech and Offensive Language Detection in Bengali (2022.aacl-main)
Copied to clipboard
| Challenge: | Existing research on hate speech detection in English does not cover low-resource languages like Bengali. |
| Approach: | They develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets. |
| Outcome: | The proposed model outperforms other models on training actual and romanized datasets by interpreting the semantic expressions better. |