Papers by Jakub Piskorski

10 papers
Multilingual Multifaceted Understanding of Online News in Terms of Genre, Framing, and Persuasion Techniques (2023.acl-long)

Copied to clipboard

Challenge: a new dataset of news articles is presented that covers genre, framing, and persuasion techniques.
Approach: They propose a multilingual multifacet dataset of news articles annotated for genre, framing and persuasion techniques.
Outcome: The proposed dataset contains 1,612 news articles covering recent news on current topics of public interest in six European languages.
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)

Copied to clipboard

Challenge: This work presents a corpus manually annotated with named entities for six Slavic languages .
Approach: They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models .
Outcome: The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier.
Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation Campaign (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper we examine the complexity of persuasion technique annotation in a multilingual annotation campaign involving 6 languages and approximately 40 annotators.
Approach: They propose a word embedding-based annotator agreement metric and propose 'holistic IAA' metric to measure the coherence of the entire dataset.
Outcome: The proposed method is compared with the existing IAA metrics and its correlation with the results.
Entity Framing and Role Portrayal in the News (2025.findings-acl)

Copied to clipboard

Challenge: a dataset of news articles containing 22 fine-grained characters is annotated for entity framing and role portrayal . the dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change .
Approach: They propose a multilingual and hierarchical corpus annotated for entity framing and role portrayal in news articles.
Outcome: The proposed dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change . the authors report evaluation results on state-of-the-art multilingual transformers and hierarchical zero-shot learning using LLMs at the level of a document, paragraph, and sentence .
PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News Articles (2025.acl-long)

Copied to clipboard

Challenge: a new dataset of news articles annotated for narratives provides a framework for narrative detection . recurring narratives can propagate with very high velocity across audiences, languages and countries .
Approach: They propose a multilingual dataset annotated for narratives using two-level taxonomies . they define narrative as a recurring, repetitive, overt or implicit claim that promotes a specific interpretation or viewpoint on an ongoing topic .
Outcome: The proposed dataset will foster research in narrative detection and enable new research directions . the authors identify multiple narratives in the same article, and the results are published online .
Exploring the Usability of Persuasion Techniques for Downstream Misinformation-related Classification Tasks (2024.lrec-main)

Copied to clipboard

Challenge: systematically explore the predictive power of features derived from Persuasion Techniques detected in texts for different tasks of interest for media analysis.
Approach: They propose a set of meaningful features aimed at capturing persuasiveness of a text . they also assess the discriminatory power of these features in different text classification tasks .
Outcome: The proposed features can be applied to detecting mis/disinformation, fake news, propaganda, partisan news and conspiracy theories.
NarratEX Dataset: Explaining the Dominant Narratives in News Texts (2025.findings-emnlp)

Copied to clipboard

Challenge: a dataset is created to explain the choice of the dominant narrative in a news article . the dataset is intended to address discourse polarization and propaganda detection .
Approach: They propose a dataset for explaining the choice of the dominant narrative in a news article . the dataset is annotated manually with a dominant narrative and sub-narrative labels .
Outcome: The proposed dataset is designed to explain the choice of the dominant narrative in a news article.
Extraction of Information Provision Activity Requirements from EU Acquis (2025.emnlp-industry)

Copied to clipboard

Challenge: Using knowledge-, classical ML-, transformer-, and generative AI-based approaches, we extract structured information from EU acquis documents.
Approach: They propose a task of Information Provision Activity Requirement Extraction to identify text fragments that introduce an obligation to provide information and the extraction of structured information about the key entities involved.
Outcome: The proposed task is based on knowledge-, classical ML-, transformer-, and generative AI-based approaches.
Resources and Experiments on Sentiment Classification for Georgian (2022.lrec-1)

Copied to clipboard

Challenge: a dataset for sentiment classification and semantic polarity dictionary for Georgian is available . a large number of linguistic resources are available for sentiment analysis for this language .
Approach: They propose to create the first publicly available annotated dataset for sentiment classification and semantic polarity dictionary for Georgian.
Outcome: The results are on par with state-of-the-art models for well-studied languages . the authors compare knowledge-and machine learning-based models to a well-supported language .
New Benchmark Corpus and Models for Fine-grained Event Classification: To BERT or not to BERT? (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for fine-grained event classification are of tiny size, ranging from 5-10K events.
Approach: They propose to use ACLED data for fine-grained event classification . they compare performance of various state-of-the-art models on these datasets .
Outcome: The proposed models perform better on micro (94.3-94.9%) and macro F1 (86.0-88.9%) the proposed models are robust and the performance is dependent on training data size.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations