Papers by Michał Marcińczuk
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | This work presents a corpus manually annotated with named entities for six Slavic languages . |
| Approach: | They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models . |
| Outcome: | The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier. |
PST 2.0 – Corpus of Polish Spatial Texts (2020.lrec-1)
Copied to clipboard
| Challenge: | In this paper, we focus on modeling spatial expressions in texts. |
| Approach: | They propose guidelines for annotating the PST 2.0 corpus of Polish Spatial Texts based on existing standards for English and discuss modifications to the guidelines to the characteristics of the language. |
| Outcome: | The proposed framework is based on three existing standards for English and ISO-Space1.4 from SpaceEval 2014 . |