| Challenge: | In this paper, we focus on modeling spatial expressions in texts. |
| Approach: | They propose guidelines for annotating the PST 2.0 corpus of Polish Spatial Texts based on existing standards for English and discuss modifications to the guidelines to the characteristics of the language. |
| Outcome: | The proposed framework is based on three existing standards for English and ISO-Space1.4 from SpaceEval 2014 . |
Similar Papers
Polish Corpus of Annotated Descriptions of Images (L18-1)
Copied to clipboard
| Challenge: | a new dataset of image descriptions is presented in Polish . the dataset is too small for training a sophisticated language-vision system. |
| Approach: | They propose to use a Polish dataset to analyze image descriptions . the descriptions are morphosyntactically analysed and annotated by human annotators . |
| Outcome: | The proposed model learns about the inter-modal correspondences between language and vision. |
Manually Annotated Corpus of Polish Texts Published between 1830 and 1918 (L18-1)
Copied to clipboard
| Challenge: | a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
| Approach: | The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
| Outcome: | The corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser Development (2024.lrec-main)
Copied to clipboard
Maciej Ogrodniczuk, Aleksandra Tomaszewska, Daniel Ziembicki, Sebastian Żurowski, Ryszard Tuora, Aleksandra Zwierzchowska
| Challenge: | The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation. |
| Approach: | They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework. |
| Outcome: | The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators. |
Annotation of metaphorical expressions in the Basic Corpus of Polish Metaphors (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Polish texts annotated with metaphorical expressions is composed of two parts of comparable size, selected from two subcorpora of the Polish National Corpus . |
| Approach: | They propose to use a procedure to annotate metaphorical expressions in Polish texts using two different subcorpora of the Polish National Corpus . they propose several features to classify metaphorical Expressions identified in texts. |
| Outcome: | The proposed procedure is based on the MIPVU procedure and focuses on neologistic derivatives that have metaphorical properties. |
Interannotator Agreement for Lexico-Semantic Annotation of a Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a method for lexico-semantic annotation of the Basic Corpus of Polish Metaphors is described . the procedure is composed of three steps: deciding whether a particular occurrence of a word is asemantics or strictly grammatical. |
| Approach: | They propose a procedure for lexico-semantic annotation of the Basic Corpus of Polish Metaphor . procedure corrects morphosyntactic annotation of part of corpus that is automatically annotated . |
| Outcome: | The proposed procedure corrects the morphosyntactic annotation of part of the corpus . it is composed of three steps: deciding whether a word is asemantic or strictly grammatical . preliminary results show that the procedure is adequate for the task . |
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)
Copied to clipboard
| Challenge: | a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer. |
| Approach: | They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph. |
| Outcome: | The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees. |
DiaBiz.Kom - towards a Polish Dialogue Act Corpus Based on ISO 24617-2 Standard (2022.coling-1)
Copied to clipboard
Marcin Oleksy, Jan Wieczorek, Dorota Drużyłowska, Julia Klyus, Aleksandra Domogała, Krzysztof Hwaszcz, Hanna Kędzierska, Daria Mikoś, Anita Wróż
| Challenge: | DiaBiz.Kom is the first corpus of dialogue texts in the Polish language . it contains transcriptions of telephone conversations conducted according to a prepared scenario. |
| Approach: | They describe the specification and evaluation of DiaBiz.Kom - the corpus of dialogue texts in Polish. |
| Outcome: | The proposed corpus contains transcriptions of telephone conversations conducted according to a prepared scenario and will be used to develop a system of dialog analysis and modules for creating advanced chatbots. |
Czech Text Document Corpus v 2.0 (L18-1)
Copied to clipboard
| Challenge: | a corpus of text documents for automatic document classification in Czech is presented . paper aims to facilitate a straightforward comparison of document classification approaches on Czech data . |
| Approach: | This paper introduces a collection of text documents for automatic document classification in Czech language. |
| Outcome: | The proposed corpus is based on the Czech news agency's real newspaper articles . it is used for evaluation of multi-label document classification approaches . |
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)
Copied to clipboard
Špela Arhar Holdt, Jaka Čibej, Kaja Dobrovoljc, Tomaž Erjavec, Polona Gantar, Simon Krek, Tina Munda, Nejc Robida, Luka Terčon, Slavko Zitnik
| Challenge: | a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years. |
| Approach: | They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene. |
| Outcome: | The revised corpus, built on its predecessor, doubles in size and depth of annotation layers. |
The Potsdam Commentary Corpus 2.2: Extending Annotations for Shallow Discourse Parsing (2020.lrec-1)
Copied to clipboard
| Challenge: | Potsdam Commentary Corpus 2.2 is a german corpus of news editorials annotated on several levels. |
| Approach: | They propose to add relation senses to an already existing layer of discourse connectives and their arguments and a new layer with additional coherence relation types to the potsdam commentary corpus. |
| Outcome: | The proposed corpus is more usable for shallow discourse parsing. |