Challenge: a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Approach: The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Outcome: The corpus provides three layers: transliteration, transcription and morphosyntactic annotation.

Similar Papers

Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser Development (2024.lrec-main)

Copied to clipboard

Challenge: The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation.
Approach: They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework.
Outcome: The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators.
Polish Corpus of Annotated Descriptions of Images (L18-1)

Copied to clipboard

Challenge: a new dataset of image descriptions is presented in Polish . the dataset is too small for training a sophisticated language-vision system.
Approach: They propose to use a Polish dataset to analyze image descriptions . the descriptions are morphosyntactically analysed and annotated by human annotators .
Outcome: The proposed model learns about the inter-modal correspondences between language and vision.
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)

Copied to clipboard

Challenge: a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer.
Approach: They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph.
Outcome: The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees.
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
Approach: They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems.
Outcome: The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
DiaBiz – an Annotated Corpus of Polish Call Center Dialogs (2022.lrec-1)

Copied to clipboard

Challenge: DiaBiz is a large corpus of phone conversations from different business domains . it contains nearly 410 hours of recordings and over 3 million words of transcribed speech.
Approach: They introduce DiaBiz, a large, annotated, multimodal corpus of Polish telephone conversations . it is a multimodal, multi-modal corpor of 4036 phone conversations from nine different domains .
Outcome: The corpus of 4036 phone conversations in Poland is 410 hours long and contains over 3 million words of transcribed speech.
Interannotator Agreement for Lexico-Semantic Annotation of a Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a method for lexico-semantic annotation of the Basic Corpus of Polish Metaphors is described . the procedure is composed of three steps: deciding whether a particular occurrence of a word is asemantics or strictly grammatical.
Approach: They propose a procedure for lexico-semantic annotation of the Basic Corpus of Polish Metaphor . procedure corrects morphosyntactic annotation of part of corpus that is automatically annotated .
Outcome: The proposed procedure corrects the morphosyntactic annotation of part of the corpus . it is composed of three steps: deciding whether a word is asemantic or strictly grammatical . preliminary results show that the procedure is adequate for the task .
An Application for Building a Polish Telephone Speech Corpus (L18-1)

Copied to clipboard

Challenge: Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance.
Approach: They propose to build a tool for speech corpus collection of a specific domain content.
Outcome: The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks.
A Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper Collection (2024.lrec-main)

Copied to clipboard

Challenge: a curated corpus of Slovenian historical newspapers is a complex undertaking requiring multiple steps to prepare . a shoestring budget is required to produce a corpus that is billion-words in size .
Approach: They propose a lightweight approach to producing high-quality corpora using OCR . they use noisy OCR-ed data from the National and University Library of Slovenia .
Outcome: The proposed method produces a billion-word giga-corpus of Slovenian historical newspapers from the 18th, 19th and 20th centuries on a shoestring budget.
Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles (2024.lrec-main)

Copied to clipboard

Challenge: Literature corpus building is relatively nascent, and standardized procedures for curating literary corpora are not yet developed.
Approach: They propose a workflow for the creation and reuse of literary corpora using a metadata-enriched Polish Novel Corpus from the 19th and 20th centuries.
Outcome: The proposed workflow includes a multi-stage metadata enrichment and verification process and efficient data collection and data sharing according to the FAIR principles and 5- and 7-star data standards.
Corpus REDEWIEDERGABE (2020.lrec-1)

Copied to clipboard

Challenge: The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind.
Approach: This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR).
Outcome: The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations