Challenge: Social media platforms provide an ideal environment to spread misinformation, where social bots can accelerate the spread.
Approach: They construct a large-scale dataset that includes annotations for misinformation and social bots on the Sina Weibo platform.
Outcome: The proposed dataset contains 65,749 social bots and 345,886 genuine accounts, annotated using a weakly supervised annotator.

Similar Papers

Dynamic Simulation Framework for Disinformation Dissemination and Correction With Social Bots (2025.findings-emnlp)

Copied to clipboard

Challenge: Current studies rely on simplistic user and network modeling and neglect dynamic behavior of bots.
Approach: They propose a multi-agent-based framework for disinformation dissemination . it incorporates both malicious and legitimate bots and allows quantitative evaluation of correction strategies.
Outcome: The proposed framework incorporates both malicious and legitimate bots and their controlled dynamic participation allows for quantitative analysis of correction strategies.
MisinfoEval: Generative AI in the Era of “Alternative Facts” (2024.emnlp-main)

Copied to clipboard

Challenge: Existing efforts to address misinformation on social media platforms are hampered by user biases and scalability challenges.
Approach: They propose a framework for generating and comprehensively evaluating large language model based misinformation interventions using a simulated social media environment and personalized explanations tailored to users' beliefs.
Outcome: The proposed framework improves accuracy at reliability labeling by up to 41.72% and personalized explanations appeal to users' pre-existing values.
What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection (2024.acl-long)

Copied to clipboard

Challenge: Social media bot detection has always been an arms race between advancements in machine learning and adversarial bot strategies to evade detection.
Approach: They propose a mixture-of-heterogeneous-experts framework to divide and conquer diverse user information modalities and propose LLM-guided manipulation of user textual and structured information to evade detection.
Outcome: The proposed framework outperforms state-of-the-art baselines on 1,000 annotated examples while bringing down existing detectors by 29.6% and harming calibration and reliability of bot detection systems.
Countering Misinformation via Emotional Response Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Social media platforms (SMPs) are one of the most effective ways to spread misinformation by engaging in constructive dialogue with users who spread – often in good faith – misleading messages.
Approach: They propose to use social correction to engage in constructive dialogue with users who spread misleading messages.
Outcome: The proposed dataset shows that it improves on previous studies on claim-response pairs and the author-reviewer pipeline.
Multimodal Pipeline for Collection of Misinformation Data from Telegram (2022.lrec-1)

Copied to clipboard

Challenge: a large portion of misinformation is spread via multimodal means, such as images and videos . a new pipeline for collecting misinformation from Telegram allows us to collect a greater variety of mis-information examples .
Approach: They propose to use AI to understand misinformation flow across social media platforms . they collect data from Telegram groups which promote COVID-19 misinformation .
Outcome: The proposed dataset contains almost one million messages from 2k different public channels related to spreading COVID-19 misinformation.
Tackling Fake News Detection by Continually Improving Social Context Representations using Graph Neural Networks (2022.acl-long)

Copied to clipboard

Challenge: Social media has enabled the propagation of fake news, text published by news sources with an intent to spread misinformation and sway beliefs.
Approach: They propose to use inference operators to analyze social media for fake news spread to uncover unobserved interactions between documents and users' engagement patterns.
Outcome: The proposed algorithms improve the performance of two fake news detection tasks.
Do Models of Mental Health Based on Social Media Data Generalize? (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing literature on the validity of proxy-based methods for annotating mental health status in social media has raised new concerns regarding their use in clinical applications.
Approach: They explore the generalization ability of machine learning classifiers trained to detect depression in individuals across multiple social media platforms.
Outcome: The proposed methods show that they can be used to train and analyze large datasets and that they are robust to large dataset sizes.
Rumor Detection on Social Media: Datasets, Methods and Opportunities (D19-50)

Copied to clipboard

Challenge: Social media platforms are used for information gathering, but they also lead to the spreading of rumors and fake news.
Approach: This paper presents a comprehensive list of datasets used for rumor detection . it also reviews the important studies based on what types of information they exploit .
Outcome: This paper presents an overview of the recent studies in the rumor detection field . it provides a comprehensive list of datasets used for rumour detection .
Identifying and Understanding User Reactions to Deceptive and Trusted Social News Sources (P18-2)

Copied to clipboard

Challenge: a new study examines how users react to news sources with different levels of credibility . a recent study found that 59% of bitly-URLs on Twitter are shared without ever being read .
Approach: They develop a model to classify user reactions into one of nine types . they also measure the speed and type of reaction for trusted and deceptive news sources .
Outcome: The proposed model classifies user reactions into one of nine types, such as answer, elaboration, and question, etc.
FACTOID: A New Dataset for Identifying Misinformation Spreaders and Political Bias (2022.lrec-1)

Copied to clipboard

Challenge: Proactively identifying misinformation spreaders is an important step towards mitigating the impact of fake news on our society.
Approach: They propose a new reddit dataset for fake news spreader analysis, called FACTOID, which tracks political discussions on Reddit since the beginning of 2020.
Outcome: The proposed dataset contains over 4K users with 3.4M posts and includes their credibility level (very low to very high) and political bias strength (extreme right to extreme left).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations