Manually Annotated Corpus of Polish Texts Published between 1830 and 1918 (L18-1)
Copied to clipboard
| Challenge: | a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
| Approach: | The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
| Outcome: | The corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
Similar Papers
Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser Development (2024.lrec-main)
Copied to clipboard
Maciej Ogrodniczuk, Aleksandra Tomaszewska, Daniel Ziembicki, Sebastian Żurowski, Ryszard Tuora, Aleksandra Zwierzchowska
| Challenge: | The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation. |
| Approach: | They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework. |
| Outcome: | The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators. |
Polish Corpus of Annotated Descriptions of Images (L18-1)
Copied to clipboard
| Challenge: | a new dataset of image descriptions is presented in Polish . the dataset is too small for training a sophisticated language-vision system. |
| Approach: | They propose to use a Polish dataset to analyze image descriptions . the descriptions are morphosyntactically analysed and annotated by human annotators . |
| Outcome: | The proposed model learns about the inter-modal correspondences between language and vision. |
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)
Copied to clipboard
| Challenge: | a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer. |
| Approach: | They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph. |
| Outcome: | The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees. |
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)
Copied to clipboard
| Challenge: | a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
| Approach: | They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems. |
| Outcome: | The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
DiaBiz – an Annotated Corpus of Polish Call Center Dialogs (2022.lrec-1)
Copied to clipboard
Piotr Pęzik, Gosia Krawentek, Sylwia Karasińska, Paweł Wilk, Paulina Rybińska, Anna Cichosz, Angelika Peljak-Łapińska, Mikołaj Deckert, Michał Adamczyk
| Challenge: | DiaBiz is a large corpus of phone conversations from different business domains . it contains nearly 410 hours of recordings and over 3 million words of transcribed speech. |
| Approach: | They introduce DiaBiz, a large, annotated, multimodal corpus of Polish telephone conversations . it is a multimodal, multi-modal corpor of 4036 phone conversations from nine different domains . |
| Outcome: | The corpus of 4036 phone conversations in Poland is 410 hours long and contains over 3 million words of transcribed speech. |
Interannotator Agreement for Lexico-Semantic Annotation of a Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a method for lexico-semantic annotation of the Basic Corpus of Polish Metaphors is described . the procedure is composed of three steps: deciding whether a particular occurrence of a word is asemantics or strictly grammatical. |
| Approach: | They propose a procedure for lexico-semantic annotation of the Basic Corpus of Polish Metaphor . procedure corrects morphosyntactic annotation of part of corpus that is automatically annotated . |
| Outcome: | The proposed procedure corrects the morphosyntactic annotation of part of the corpus . it is composed of three steps: deciding whether a word is asemantic or strictly grammatical . preliminary results show that the procedure is adequate for the task . |
An Application for Building a Polish Telephone Speech Corpus (L18-1)
Copied to clipboard
| Challenge: | Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance. |
| Approach: | They propose to build a tool for speech corpus collection of a specific domain content. |
| Outcome: | The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks. |
A Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper Collection (2024.lrec-main)
Copied to clipboard
| Challenge: | a curated corpus of Slovenian historical newspapers is a complex undertaking requiring multiple steps to prepare . a shoestring budget is required to produce a corpus that is billion-words in size . |
| Approach: | They propose a lightweight approach to producing high-quality corpora using OCR . they use noisy OCR-ed data from the National and University Library of Slovenia . |
| Outcome: | The proposed method produces a billion-word giga-corpus of Slovenian historical newspapers from the 18th, 19th and 20th centuries on a shoestring budget. |
Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles (2024.lrec-main)
Copied to clipboard
| Challenge: | Literature corpus building is relatively nascent, and standardized procedures for curating literary corpora are not yet developed. |
| Approach: | They propose a workflow for the creation and reuse of literary corpora using a metadata-enriched Polish Novel Corpus from the 19th and 20th centuries. |
| Outcome: | The proposed workflow includes a multi-stage metadata enrichment and verification process and efficient data collection and data sharing according to the FAIR principles and 5- and 7-star data standards. |
Corpus REDEWIEDERGABE (2020.lrec-1)
Copied to clipboard
| Challenge: | The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind. |
| Approach: | This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR). |
| Outcome: | The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind. |