Papers by Michał Marcińczuk

2 papers
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)

Copied to clipboard

Challenge: This work presents a corpus manually annotated with named entities for six Slavic languages .
Approach: They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models .
Outcome: The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier.
PST 2.0 – Corpus of Polish Spatial Texts (2020.lrec-1)

Copied to clipboard

Challenge: In this paper, we focus on modeling spatial expressions in texts.
Approach: They propose guidelines for annotating the PST 2.0 corpus of Polish Spatial Texts based on existing standards for English and discuss modifications to the guidelines to the characteristics of the language.
Outcome: The proposed framework is based on three existing standards for English and ISO-Space1.4 from SpaceEval 2014 .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations