Challenge: Using a unique methodology, we annotated disinformation in Polish with multiple labels indicating both intents and manipulation techniques employed.
Approach: They present a novel corpus of 15,356 Polish web articles annotated with multiple labels indicating both disinformation creators’ intents and manipulation techniques employed.
Outcome: The proposed dataset sheds light on the authors' intention and manipulation techniques in disinformation.

Similar Papers

ZenPropaganda: A Comprehensive Study on Identifying Propaganda Techniques in Russian Coronavirus-Related Media (2024.lrec-main)

Copied to clipboard

Challenge: a new classification scheme for automatic detection of propaganda techniques is proposed . the capabilities of algorithms increase the risks of propaganda impact on the audience .
Approach: They propose a novel multi-level classification scheme for automatic detection of propaganda techniques.
Outcome: The proposed classification scheme outperforms existing methods in a Russian dataset and provides a valuable resource for future research.
GerDISDETECT: A German Multilabel Dataset for Disinformation Detection (2024.lrec-main)

Copied to clipboard

Challenge: Disinformation datasets are sparse and expensive to train . annotated datasets often have only binary or multiclass labels .
Approach: They propose to use a textual dataset to detect disinformation in German . the dataset contains 39 multilabel classes with 5 top-level categories .
Outcome: The proposed dataset provides comprehensive insights into disinformation in German using a taxonomy guided annotation scheme.
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns (2024.acl-srw)

Copied to clipboard

Challenge: Existing methods for multilingual framing differ from those used in English-speaking world . framers often use loaded vocabularies to create political images or favor a particular point of view .
Approach: They use eight years of Russian-backed disinformation campaigns to examine framing . they find that disinformation campaign consistently favors specific framers .
Outcome: The proposed method underperforms and shows high disagreements in Russian-language articles . the proposed method is based on eight years of Russian-backed disinformation campaigns .
Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation (2021.acl-long)

Copied to clipboard

Challenge: Edited media frames are structured annotations with respect to intents, emotional reactions, attacks on individuals, and the implications of disinformation.
Approach: They propose a new formalism to understand visual media manipulation as structured annotations with respect to intents, emotional reactions, attacks on individuals, and the implications of disinformation.
Outcome: The proposed model obtains promising results on a dataset with 56k question-answer pairs written in rich natural language.
MALicious INTent Dataset and Inoculating LLMs for Enhanced Disinformation Detection (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies on intentionality behind disinformation do not address intent behind disinformative agents.
Approach: They propose an intent-augmented reasoning system that integrates intent analysis to mitigate the persuasive impact of disinformation.
Outcome: The proposed corpus is the first human-annotated English corpus to capture disinformation and its malicious intent.
DeFaktS: A German Dataset for Fine-Grained Disinformation Detection through Social Media Framing (2024.lrec-main)

Copied to clipboard

Challenge: Distinctively curated across various news topics, DeFaktS offers an unparalleled insight into disinformation’s diverse characteristics.
Approach: They propose to annotate every structural component and semantic element of a news piece, eliminating the need for external knowledge sources.
Outcome: The proposed dataset contains 105,855 posts with 20,008 meticulously labeled tweets and eliminates the need for external knowledge sources.
BAN-PL: A Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl Web Service (2024.lrec-main)

Copied to clipboard

Challenge: a new dataset of offensive social media content for the Polish language is presented to address this gap . access to accurate and non-synthetic datasets of social media is limited for low-resource languages .
Approach: They present a new open dataset of offensive social media content for the Polish language . authors propose to make the dataset publicly available to improve access .
Outcome: The proposed dataset includes 691,662 posts and comments from the Polish Reddit . the authors describe the dataset and apply it to real-life content moderation processes .
Rhetorical Structure Approach for Online Deception Detection: A Survey (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on how people use language to inform and misinform are relevant.
Approach: They analyze how discourse structure is applied to fake news detection on the web and social media.
Outcome: The proposed framework is applied to fake news and fake reviews detection on the web and social media.
PropaInsight: Toward Deeper Understanding of Propaganda in Terms of Techniques, Appeals, and Intent (2025.coling-main)

Copied to clipboard

Challenge: Existing research on propaganda detection does not capture the motives behind the content or its broader impact.
Approach: They propose a framework that dissects propaganda into techniques, arousal appeals, and underlying intent.
Outcome: The proposed framework improves performance in a wide range of scenarios and can be used to identify and categorize propaganda techniques.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations