Challenge: Proactively identifying misinformation spreaders is an important step towards mitigating the impact of fake news on our society.
Approach: They propose a new reddit dataset for fake news spreader analysis, called FACTOID, which tracks political discussions on Reddit since the beginning of 2020.
Outcome: The proposed dataset contains over 4K users with 3.4M posts and includes their credibility level (very low to very high) and political bias strength (extreme right to extreme left).

Similar Papers

Evons: A Dataset for Fake and Real News Virality Analysis and Prediction (2022.coling-1)

Copied to clipboard

Challenge: Existing collections of fake news articles contain claims fact-checked for veracity . Developing fake news detection models requires annotated collections of real and fake news.
Approach: They propose to annotate news articles originating from fake and real news media sources for the analysis and prediction of news virality.
Outcome: The proposed collection is compared with existing datasets which contain claims or headline and body text but can't be used for predicting fake news virality.
Analysing State-Backed Propaganda Websites: a New Dataset and Linguistic Study (2023.emnlp-main)

Copied to clipboard

Challenge: a network of doppelganger websites (impersonating genuine news sites) was discovered in 2022 . a novel dataset enables studies of disinformation networks and the training of NLP tools for disinformation detection.
Approach: They analyze two hitherto unstudied sites sharing state-backed disinformation . they perform cross-site topic clustering and perform linguistic and temporal analysis .
Outcome: The proposed dataset includes 14,053 articles, annotated with each language version, and additional metadata such as links and images.
A Survey on Predicting the Factuality and the Bias of News Media (2024.findings-acl)

Copied to clipboard

Challenge: a growing number of scholars are profiling entire news outlets to profile fake content . political bias detection is also an important topic, but the two problems have been addressed separately .
Approach: They argue that media profiling should be based on factuality and bias together . they argue that it is difficult to fact-check every single suspicious claim or article manually .
Outcome: The present level of proliferation of fake, biased, and propagandistic content online has made it impossible to fact-check every single suspicious claim or article, either manually or automatically.
Demystifying Neural Fake News via Linguistic Feature-Based Interpretation (2022.coling-1)

Copied to clipboard

Challenge: Recent advances to neural fake news generators have made it difficult to understand how misinformation generated by these models may best be confronted.
Approach: They conduct feature-based analysis to gain an interpretative understanding of the linguistic attributes that neural fake news generators may most effectively exploit.
Outcome: The proposed models are compared with models trained on subsets of features and confronted with increasingly advanced neural fake news.
Identifying and Understanding User Reactions to Deceptive and Trusted Social News Sources (P18-2)

Copied to clipboard

Challenge: a new study examines how users react to news sources with different levels of credibility . a recent study found that 59% of bitly-URLs on Twitter are shared without ever being read .
Approach: They develop a model to classify user reactions into one of nine types . they also measure the speed and type of reaction for trusted and deceptive news sources .
Outcome: The proposed model classifies user reactions into one of nine types, such as answer, elaboration, and question, etc.
Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News Detection (2020.lrec-1)

Copied to clipboard

Challenge: Prior fake news datasets lack multimodal text and image data, metadata, comment data, and fine-grained classification at the scale and breadth of their datasets.
Approach: They propose to use a multimodal dataset to build a machine learning classification model that uses text and image data to classify fake news.
Outcome: The proposed model is based on a multimodal dataset consisting of over 1 million samples from multiple categories of fake news.
BREAKING! Presenting Fake News Corpus for Automated Fact Checking (P19-2)

Copied to clipboard

Challenge: a new study shows that fake news spreads faster than mainstream articles on the same topic . however, there is no dataset containing compelling fake and questionable news articles .
Approach: They introduce manually verified corpus of compelling fake and questionable news articles on the USA politics . they plan to extend the corpus in the future and use it for automated fake news detection.
Outcome: The proposed model is based on linguistic features and will be extended in the future . it will be used to improve the existing model and improve the tools in the field of fake news detection .
Annotating and Analyzing Biased Sentences in News Articles using Crowdsourcing (2020.lrec-1)

Copied to clipboard

Challenge: a lack of publicly available news bias datasets has hindered efforts to detect subtle biases in news articles.
Approach: They propose a news bias dataset which contains sentences with bias labels . they propose to use the dataset to develop and evaluate methods for detecting news bias .
Outcome: The proposed dataset can be used for analyzing news bias and for developing and evaluating methods for news bias detection.
How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysis (2025.emnlp-main)

Copied to clipboard

Challenge: Social media platforms provide an ideal environment to spread misinformation, where social bots can accelerate the spread.
Approach: They construct a large-scale dataset that includes annotations for misinformation and social bots on the Sina Weibo platform.
Outcome: The proposed dataset contains 65,749 social bots and 345,886 genuine accounts, annotated using a weakly supervised annotator.
Words are the Window to the Soul: Language-based User Representations for Fake News Detection (2020.coling-main)

Copied to clipboard

Challenge: Existing studies on fake news classification focus on textual content, but also social context in which news are consumed.
Approach: They propose a model that creates representations of individuals on social media based only on the language they produce and uses them to detect fake news.
Outcome: The proposed model exploits the relationship between language use and connections in the social graph to assess the presence of the Echo Chamber effect in the data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations