Assessing Post Deletion in Sina Weibo: Multi-modal Classification of Hot Topics (D19-50)
Copied to clipboard
Meisam Navaki Arefi, Rajkumar Pandi, Michael Carl Tschantz, Jedidiah R. Crandall, King-wa Fu, Dahlia Qiu Shi, Miao Sha
| Challenge: | Weibo monitors and deletes posts to conform to government requirements . a recent study found that sentiment is the only indicator of censorship that is consistent across topics . |
| Approach: | They analyze a dataset of censored and uncensore censors on Weibo . they use deep learning, CNN localization, and NLP techniques to analyze the data . |
| Outcome: | The proposed analysis of censored and uncensoreded posts in Weibo shows that sentiment is the only indicator of a topic's censorship . censors can remove posts that are considered sensitive in three hours on average . |
Similar Papers
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing knowledge on how and why NLP methods make content moderation decisions is limited . authors examine how and when to use LLMs in content modeation . |
| Approach: | They use Shapley values and LLM-guided explanations to reverse-engineer content moderation decisions across countries. |
| Outcome: | The proposed methods show that they reverse-engineer content moderation decisions across countries and over time. |
SCCD: A Session-based Dataset for Chinese Cyberbullying Detection (2025.coling-main)
Copied to clipboard
| Challenge: | Existing work on cyberbullying detection in Chinese is underdeveloped due to the lack of comprehensive and reliable datasets. |
| Approach: | They propose to use Chinese social media sessions to analyze Chinese cyberbullying content to improve the quality of annotations. |
| Outcome: | The proposed dataset shows that it performs better than existing methods on Weibo and a major social media platform. |
Hashtags, Emotions, and Comments: A Large-Scale Dataset to Understand Fine-Grained Social Emotions to Online Topics (2020.emnlp-main)
Copied to clipboard
| Challenge: | A large-scale dataset is collected from Chinese microblog Sina Weibo with over 13 thousand trending topics, emotion votes in 24 fine-grained types from massive participants, and user comments to allow context understanding. |
| Approach: | They use a large-scale dataset from Chinese microblog Sina Weibo to examine readers' responses to online discussion topics. |
| Outcome: | The proposed model outperforms the human model in predicting social emotions in a multilabel classification setting. |
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings (2025.findings-acl)
Copied to clipboard
Shujian Yang, Shiyao Cui, Chuanrui Hu, Haicheng Wang, Tianwei Zhang, Minlie Huang, Jialiang Lu, Han Qiu
| Challenge: | Recent studies show that character substitutions in toxic Chinese text can confuse state-of-the-art LLMs. |
| Approach: | They propose a taxonomy of 3 perturbation strategies and 8 specific approaches in Chinese text to assess if they can detect perturbed Chinese toxic contents. |
| Outcome: | The proposed model can detect perturbed Chinese text with 8 different approaches . the proposed model is compared with 9 other LLMs from the US and China . |
Classifying Social Media Users before and after Depression Diagnosis via Their Language Usage: A Dataset and Study (2024.lrec-main)
Copied to clipboard
| Challenge: | Mental illness can negatively impact individuals’ quality of life as it is considered one of the causes of years lived with disability and it is related to high suicide rates. |
| Approach: | They collect first dataset of textual posts by same users before and after being diagnosed with depression and build multiple predictive models based on Transformers and BERT. |
| Outcome: | The proposed model can be used to detect depression and suicidal thoughts in users who are not diagnosed with depression or suicide. |
Towards Exploiting Sticker for Multimodal Sentiment Analysis in Social Media: A New Dataset and Baseline (2022.coling-1)
Copied to clipboard
| Challenge: | Sentiment analysis in social media is challenging because of the lack of context. |
| Approach: | They propose to use stickers to perform a multimodal sentiment analysis task using Chinese stickers. |
| Outcome: | The proposed model performs best compared with other models. |
Detecting Community Sensitive Norm Violations in Online Conversations (2021.findings-emnlp)
Copied to clipboard
Chan Young Park, Julia Mendelsohn, Karthik Radhakrishnan, Kinjal Jain, Tushar Kanakagiri, David Jurgens, Yulia Tsvetkov
| Challenge: | Existing efforts to identify unacceptable behavior have focused on toxicity as the sole form of community norm violation. |
| Approach: | They propose a dataset that focuses on a more complete spectrum of community norms and their violations in local conversational and global contexts. |
| Outcome: | The proposed model improves the detection of community norm violations in local conversational and global contexts. |
RESEMO: A Benchmark Chinese Dataset for Studying Responsive Emotion from Social Media Content (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on social media text processing do not focus on responsive emotion analysis. |
| Approach: | They propose a Chinese dataset named ResEmo for responsive emotion analysis, including 3813 posts with 68,781 comments collected from Weibo, the largest social media platform in China. |
| Outcome: | The proposed dataset includes 3813 posts with 68,781 comments collected from weibo, the largest social media platform in China. |
Analyzing Polarization in Social Media: Method and Application to Tweets on 21 Mass Shootings (N19-1)
Copied to clipboard
| Challenge: | a new framework for studying political polarization in social media is needed to understand how group divisions manifest in language. |
| Approach: | They propose to cluster tweet embeddings to uncover four dimensions of political polarization in social media . their results apply existing lexical methods to analyze 4.4M tweets on 21 mass shootings . |
| Outcome: | The proposed framework generates more cohesive topics than traditional models. |
How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysis (2025.emnlp-main)
Copied to clipboard
| Challenge: | Social media platforms provide an ideal environment to spread misinformation, where social bots can accelerate the spread. |
| Approach: | They construct a large-scale dataset that includes annotations for misinformation and social bots on the Sina Weibo platform. |
| Outcome: | The proposed dataset contains 65,749 social bots and 345,886 genuine accounts, annotated using a weakly supervised annotator. |