Challenge: The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation.
Approach: They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework.
Outcome: The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators.

Similar Papers

Manually Annotated Corpus of Polish Texts Published between 1830 and 1918 (L18-1)

Copied to clipboard

Challenge: a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Approach: The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Outcome: The corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
DiaBiz.Kom - towards a Polish Dialogue Act Corpus Based on ISO 24617-2 Standard (2022.coling-1)

Copied to clipboard

Challenge: DiaBiz.Kom is the first corpus of dialogue texts in the Polish language . it contains transcriptions of telephone conversations conducted according to a prepared scenario.
Approach: They describe the specification and evaluation of DiaBiz.Kom - the corpus of dialogue texts in Polish.
Outcome: The proposed corpus contains transcriptions of telephone conversations conducted according to a prepared scenario and will be used to develop a system of dialog analysis and modules for creating advanced chatbots.
ISO-based Annotated Multilingual Parallel Corpus for Discourse Markers (2022.lrec-1)

Copied to clipboard

Challenge: Discourse markers carry information about the discourse structure and organization, and also signal local dependencies or epistemic stance of speaker.
Approach: They propose an ISO-based annotated multilingual parallel corpus for discourse markers . they propose an annotation scheme for discourse relations with a plug-in to ISO 24617-2 .
Outcome: The proposed language resource is based on an ISO-based annotated multilingual parallel corpus of discourse markers.
DiaBiz – an Annotated Corpus of Polish Call Center Dialogs (2022.lrec-1)

Copied to clipboard

Challenge: DiaBiz is a large corpus of phone conversations from different business domains . it contains nearly 410 hours of recordings and over 3 million words of transcribed speech.
Approach: They introduce DiaBiz, a large, annotated, multimodal corpus of Polish telephone conversations . it is a multimodal, multi-modal corpor of 4036 phone conversations from nine different domains .
Outcome: The corpus of 4036 phone conversations in Poland is 410 hours long and contains over 3 million words of transcribed speech.
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)

Copied to clipboard

Challenge: a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer.
Approach: They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph.
Outcome: The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees.
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.
An Application for Building a Polish Telephone Speech Corpus (L18-1)

Copied to clipboard

Challenge: Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance.
Approach: They propose to build a tool for speech corpus collection of a specific domain content.
Outcome: The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks.
Interannotator Agreement for Lexico-Semantic Annotation of a Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a method for lexico-semantic annotation of the Basic Corpus of Polish Metaphors is described . the procedure is composed of three steps: deciding whether a particular occurrence of a word is asemantics or strictly grammatical.
Approach: They propose a procedure for lexico-semantic annotation of the Basic Corpus of Polish Metaphor . procedure corrects morphosyntactic annotation of part of corpus that is automatically annotated .
Outcome: The proposed procedure corrects the morphosyntactic annotation of part of the corpus . it is composed of three steps: deciding whether a word is asemantic or strictly grammatical . preliminary results show that the procedure is adequate for the task .
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)

Copied to clipboard

Challenge: Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications.
Approach: They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages.
Outcome: The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages.
PST 2.0 – Corpus of Polish Spatial Texts (2020.lrec-1)

Copied to clipboard

Challenge: In this paper, we focus on modeling spatial expressions in texts.
Approach: They propose guidelines for annotating the PST 2.0 corpus of Polish Spatial Texts based on existing standards for English and discuss modifications to the guidelines to the characteristics of the language.
Outcome: The proposed framework is based on three existing standards for English and ISO-Space1.4 from SpaceEval 2014 .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations