Challenge: Social media has been used by troll farms to promote political agendas . trolled farms employ people to provoke conflict via the use of inflammatory or provocative comments.
Approach: They analyze the use of self-supervision with less than 100 troll accounts as training data to determine whether a trolled account is labeled as a Russian trol farm.
Outcome: The proposed methods improve classification performance by nearly 4% F1 and use self-supervision with less than 100 troll accounts as training data.

Similar Papers

Detecting Troll Tweets in a Bilingual Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a large amount of troll accounts have emerged with efforts to manipulate public opinion on social network sites . a recent study found that trolled tweets spread misinformation, fake news, and propaganda . we use supervised classification to detect trol tweets in both English and Russian .
Approach: They propose to detect troll tweets in English and Russian using machine learning algorithms . they use monolingual, cross-lingual, and bilingual training scenarios .
Outcome: The proposed method uses monolingual, cross-lingual, and bilingual training scenarios.
ELF22: A Context-based Counter Trolling Dataset to Combat Internet Trolls (2022.lrec-1)

Copied to clipboard

Challenge: a new dataset aims to automate the method to counter trolls . trolleds cause psychological damage to individuals and increase social costs .
Approach: They propose to use a dataset to generate counter responses by varying counter responses according to a given strategy.
Outcome: The proposed method improves strategy-controlled sentence generation.
Learning to love diligent trolls: Accounting for rater effects in the dialogue safety task (2023.findings-emnlp)

Copied to clipboard

Challenge: Xu et al., 2018: chatbots generate offensive utterances, which must be avoided . he proposes a solution that can learn from user interactions in a way that is robust to trolls .
Approach: They propose a method to learn from user feedback in a way that is robust to trolls . they propose multiple users rate each utterance, then perform latent class analysis to infer correct labels.
Outcome: The proposed solution can infer training labels with high accuracy when trolls are consistent, even when a majority are trolled.
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation.
Approach: They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic .
Outcome: The proposed models do not generalize, indicating heterogeneous political users.
Pioneering Bot Detection on Polish Reddit at the Comment Level (2026.eacl-srw)

Copied to clipboard

Challenge: 40,000 comments, 58% bot-comment prevalence, are used for comment-level bot detection within Polish Reddit communities.
Approach: They construct a dataset with 40,000 comments, 58% bot-comment prevalence, which provides labels for the subsequent model training.
Outcome: The proposed model trains on a Polish Reddit dataset with a linguistically mixed dataset and achieves strong performance and temporal generalization to 2025.
Modeling Trolling in Social Media Conversations (L18-1)

Copied to clipboard

Challenge: a new classification of trolling allows for comment-based analysis from both the trolls' and the responders' perspectives . a trolled's intentions may cause a negative psychological impact on the participants .
Approach: They propose a trolling categorization that allows comment-based analysis from both trolls' and responders' perspectives . they annotate and release a dataset containing excerpts of Reddit conversations involving suspected trolled users .
Outcome: The proposed model allows comment-based analysis from both the trolls' and the responders' perspectives.
Using RL to Identify Divisive Perspectives Improves LLMs Abilities to Identify Communities on Social Media (2024.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show improvements on Reddit and Twitter data .
Approach: They propose to take advantage of Large Language Models (LLMs) to better identify user communities.
Outcome: The proposed model improves on Reddit and Twitter data and tasks of community detection, bot detection, and news media profiling.
What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection (2024.acl-long)

Copied to clipboard

Challenge: Social media bot detection has always been an arms race between advancements in machine learning and adversarial bot strategies to evade detection.
Approach: They propose a mixture-of-heterogeneous-experts framework to divide and conquer diverse user information modalities and propose LLM-guided manipulation of user textual and structured information to evade detection.
Outcome: The proposed framework outperforms state-of-the-art baselines on 1,000 annotated examples while bringing down existing detectors by 29.6% and harming calibration and reliability of bot detection systems.
RuBia: A Russian Language Bias Detection Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) tend to learn the social and cultural biases present in the raw pre-training data.
Approach: They present a bias detection dataset specifically designed for the Russian language, dubbed RuBia, which is divided into 4 domains: gender, nationality, socio-economic status, and diverse.
Outcome: The proposed dataset is designed to detect bias in the Russian language and is based on 2,000 unique sentence pairs spread over 19 subdomains.
BotPercent: Estimating Bot Populations in Twitter Communities (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to bot detection are agnostic to social environments the bots operate in . however, standard approaches are not a good fit for the social environments they operate in.
Approach: They propose a method that estimates the percentage of Twitter bots given a community . they use Twitter bot detection datasets and feature-, text-, and graph-based models adjusted to a particular community based on Twitter .
Outcome: The proposed method achieves state-of-the-art in community-level Twitter bot detection across balanced and imbalanced class distribution settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations