Papers by Jakub Piskorski
Multilingual Multifaceted Understanding of Online News in Terms of Genre, Framing, and Persuasion Techniques (2023.acl-long)
Copied to clipboard
| Challenge: | a new dataset of news articles is presented that covers genre, framing, and persuasion techniques. |
| Approach: | They propose a multilingual multifacet dataset of news articles annotated for genre, framing and persuasion techniques. |
| Outcome: | The proposed dataset contains 1,612 news articles covering recent news on current topics of public interest in six European languages. |
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | This work presents a corpus manually annotated with named entities for six Slavic languages . |
| Approach: | They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models . |
| Outcome: | The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier. |
Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation Campaign (2023.emnlp-main)
Copied to clipboard
| Challenge: | In this paper we examine the complexity of persuasion technique annotation in a multilingual annotation campaign involving 6 languages and approximately 40 annotators. |
| Approach: | They propose a word embedding-based annotator agreement metric and propose 'holistic IAA' metric to measure the coherence of the entire dataset. |
| Outcome: | The proposed method is compared with the existing IAA metrics and its correlation with the results. |
Entity Framing and Role Portrayal in the News (2025.findings-acl)
Copied to clipboard
Tarek Mahmoud, Zhuohan Xie, Dimitar Iliyanov Dimitrov, Nikolaos Nikolaidis, Purificação Silvano, Roman Yangarber, Shivam Sharma, Elisa Sartori, Nicolas Stefanovitch, Giovanni Da San Martino, Jakub Piskorski, Preslav Nakov
| Challenge: | a dataset of news articles containing 22 fine-grained characters is annotated for entity framing and role portrayal . the dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change . |
| Approach: | They propose a multilingual and hierarchical corpus annotated for entity framing and role portrayal in news articles. |
| Outcome: | The proposed dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change . the authors report evaluation results on state-of-the-art multilingual transformers and hierarchical zero-shot learning using LLMs at the level of a document, paragraph, and sentence . |
PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News Articles (2025.acl-long)
Copied to clipboard
Nikolaos Nikolaidis, Nicolas Stefanovitch, Purificação Silvano, Dimitar Iliyanov Dimitrov, Roman Yangarber, Nuno Guimarães, Elisa Sartori, Ion Androutsopoulos, Preslav Nakov, Giovanni Da San Martino, Jakub Piskorski
| Challenge: | a new dataset of news articles annotated for narratives provides a framework for narrative detection . recurring narratives can propagate with very high velocity across audiences, languages and countries . |
| Approach: | They propose a multilingual dataset annotated for narratives using two-level taxonomies . they define narrative as a recurring, repetitive, overt or implicit claim that promotes a specific interpretation or viewpoint on an ongoing topic . |
| Outcome: | The proposed dataset will foster research in narrative detection and enable new research directions . the authors identify multiple narratives in the same article, and the results are published online . |
Exploring the Usability of Persuasion Techniques for Downstream Misinformation-related Classification Tasks (2024.lrec-main)
Copied to clipboard
| Challenge: | systematically explore the predictive power of features derived from Persuasion Techniques detected in texts for different tasks of interest for media analysis. |
| Approach: | They propose a set of meaningful features aimed at capturing persuasiveness of a text . they also assess the discriminatory power of these features in different text classification tasks . |
| Outcome: | The proposed features can be applied to detecting mis/disinformation, fake news, propaganda, partisan news and conspiracy theories. |
NarratEX Dataset: Explaining the Dominant Narratives in News Texts (2025.findings-emnlp)
Copied to clipboard
Nuno Guimarães, Purificação Silvano, Ricardo Campos, Alipio Jorge, Ana Filipa Pacheco, Dimitar Iliyanov Dimitrov, Nikolaos Nikolaidis, Roman Yangarber, Elisa Sartori, Nicolas Stefanovitch, Preslav Nakov, Jakub Piskorski, Giovanni Da San Martino
| Challenge: | a dataset is created to explain the choice of the dominant narrative in a news article . the dataset is intended to address discourse polarization and propaganda detection . |
| Approach: | They propose a dataset for explaining the choice of the dominant narrative in a news article . the dataset is annotated manually with a dominant narrative and sub-narrative labels . |
| Outcome: | The proposed dataset is designed to explain the choice of the dominant narrative in a news article. |
Extraction of Information Provision Activity Requirements from EU Acquis (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Using knowledge-, classical ML-, transformer-, and generative AI-based approaches, we extract structured information from EU acquis documents. |
| Approach: | They propose a task of Information Provision Activity Requirement Extraction to identify text fragments that introduce an obligation to provide information and the extraction of structured information about the key entities involved. |
| Outcome: | The proposed task is based on knowledge-, classical ML-, transformer-, and generative AI-based approaches. |
Resources and Experiments on Sentiment Classification for Georgian (2022.lrec-1)
Copied to clipboard
| Challenge: | a dataset for sentiment classification and semantic polarity dictionary for Georgian is available . a large number of linguistic resources are available for sentiment analysis for this language . |
| Approach: | They propose to create the first publicly available annotated dataset for sentiment classification and semantic polarity dictionary for Georgian. |
| Outcome: | The results are on par with state-of-the-art models for well-studied languages . the authors compare knowledge-and machine learning-based models to a well-supported language . |
New Benchmark Corpus and Models for Fine-grained Event Classification: To BERT or not to BERT? (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for fine-grained event classification are of tiny size, ranging from 5-10K events. |
| Approach: | They propose to use ACLED data for fine-grained event classification . they compare performance of various state-of-the-art models on these datasets . |
| Outcome: | The proposed models perform better on micro (94.3-94.9%) and macro F1 (86.0-88.9%) the proposed models are robust and the performance is dependent on training data size. |