How Many Users Are Enough? Exploring Semi-Supervision and Stylometric Features to Uncover a Russian Troll Farm (D19-50)
Copied to clipboard
| Challenge: | Social media has been used by troll farms to promote political agendas . trolled farms employ people to provoke conflict via the use of inflammatory or provocative comments. |
| Approach: | They analyze the use of self-supervision with less than 100 troll accounts as training data to determine whether a trolled account is labeled as a Russian trol farm. |
| Outcome: | The proposed methods improve classification performance by nearly 4% F1 and use self-supervision with less than 100 troll accounts as training data. |
Similar Papers
Detecting Troll Tweets in a Bilingual Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a large amount of troll accounts have emerged with efforts to manipulate public opinion on social network sites . a recent study found that trolled tweets spread misinformation, fake news, and propaganda . we use supervised classification to detect trol tweets in both English and Russian . |
| Approach: | They propose to detect troll tweets in English and Russian using machine learning algorithms . they use monolingual, cross-lingual, and bilingual training scenarios . |
| Outcome: | The proposed method uses monolingual, cross-lingual, and bilingual training scenarios. |
ELF22: A Context-based Counter Trolling Dataset to Combat Internet Trolls (2022.lrec-1)
Copied to clipboard
| Challenge: | a new dataset aims to automate the method to counter trolls . trolleds cause psychological damage to individuals and increase social costs . |
| Approach: | They propose to use a dataset to generate counter responses by varying counter responses according to a given strategy. |
| Outcome: | The proposed method improves strategy-controlled sentence generation. |
Learning to love diligent trolls: Accounting for rater effects in the dialogue safety task (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Xu et al., 2018: chatbots generate offensive utterances, which must be avoided . he proposes a solution that can learn from user interactions in a way that is robust to trolls . |
| Approach: | They propose a method to learn from user feedback in a way that is robust to trolls . they propose multiple users rate each utterance, then perform latent class analysis to infer correct labels. |
| Outcome: | The proposed solution can infer training labels with high accuracy when trolls are consistent, even when a majority are trolled. |
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation. |
| Approach: | They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic . |
| Outcome: | The proposed models do not generalize, indicating heterogeneous political users. |
Pioneering Bot Detection on Polish Reddit at the Comment Level (2026.eacl-srw)
Copied to clipboard
| Challenge: | 40,000 comments, 58% bot-comment prevalence, are used for comment-level bot detection within Polish Reddit communities. |
| Approach: | They construct a dataset with 40,000 comments, 58% bot-comment prevalence, which provides labels for the subsequent model training. |
| Outcome: | The proposed model trains on a Polish Reddit dataset with a linguistically mixed dataset and achieves strong performance and temporal generalization to 2025. |
Modeling Trolling in Social Media Conversations (L18-1)
Copied to clipboard
| Challenge: | a new classification of trolling allows for comment-based analysis from both the trolls' and the responders' perspectives . a trolled's intentions may cause a negative psychological impact on the participants . |
| Approach: | They propose a trolling categorization that allows comment-based analysis from both trolls' and responders' perspectives . they annotate and release a dataset containing excerpts of Reddit conversations involving suspected trolled users . |
| Outcome: | The proposed model allows comment-based analysis from both the trolls' and the responders' perspectives. |
Using RL to Identify Divisive Perspectives Improves LLMs Abilities to Identify Communities on Social Media (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Experimental results show improvements on Reddit and Twitter data . |
| Approach: | They propose to take advantage of Large Language Models (LLMs) to better identify user communities. |
| Outcome: | The proposed model improves on Reddit and Twitter data and tasks of community detection, bot detection, and news media profiling. |
What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection (2024.acl-long)
Copied to clipboard
| Challenge: | Social media bot detection has always been an arms race between advancements in machine learning and adversarial bot strategies to evade detection. |
| Approach: | They propose a mixture-of-heterogeneous-experts framework to divide and conquer diverse user information modalities and propose LLM-guided manipulation of user textual and structured information to evade detection. |
| Outcome: | The proposed framework outperforms state-of-the-art baselines on 1,000 annotated examples while bringing down existing detectors by 29.6% and harming calibration and reliability of bot detection systems. |
RuBia: A Russian Language Bias Detection Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) tend to learn the social and cultural biases present in the raw pre-training data. |
| Approach: | They present a bias detection dataset specifically designed for the Russian language, dubbed RuBia, which is divided into 4 domains: gender, nationality, socio-economic status, and diverse. |
| Outcome: | The proposed dataset is designed to detect bias in the Russian language and is based on 2,000 unique sentence pairs spread over 19 subdomains. |
BotPercent: Estimating Bot Populations in Twitter Communities (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to bot detection are agnostic to social environments the bots operate in . however, standard approaches are not a good fit for the social environments they operate in. |
| Approach: | They propose a method that estimates the percentage of Twitter bots given a community . they use Twitter bot detection datasets and feature-, text-, and graph-based models adjusted to a particular community based on Twitter . |
| Outcome: | The proposed method achieves state-of-the-art in community-level Twitter bot detection across balanced and imbalanced class distribution settings. |