An Application for Building a Polish Telephone Speech Corpus (L18-1)

Copied to clipboard

Challenge: Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance.
Approach: They propose to build a tool for speech corpus collection of a specific domain content.
Outcome: The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks.

Similar Papers

DiaBiz – an Annotated Corpus of Polish Call Center Dialogs (2022.lrec-1)

Copied to clipboard

Challenge: DiaBiz is a large corpus of phone conversations from different business domains . it contains nearly 410 hours of recordings and over 3 million words of transcribed speech.
Approach: They introduce DiaBiz, a large, annotated, multimodal corpus of Polish telephone conversations . it is a multimodal, multi-modal corpor of 4036 phone conversations from nine different domains .
Outcome: The corpus of 4036 phone conversations in Poland is 410 hours long and contains over 3 million words of transcribed speech.
Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech (2024.lrec-main)

Copied to clipboard

Challenge: The Dutch Dialect Database contains dialectal variations of Dutch recorded in the second half of the twentieth century.
Approach: They propose to create a corpus containing audio recordings and orthographic transcriptions of Dutch dialects recorded in the second half of the 20th century.
Outcome: The Dutch Dialect Database contains dialectal variations recorded all over the Netherlands in the second half of the twentieth century.
Manually Annotated Corpus of Polish Texts Published between 1830 and 1918 (L18-1)

Copied to clipboard

Challenge: a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Approach: The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Outcome: The corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)

Copied to clipboard

Challenge: a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer.
Approach: They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph.
Outcome: The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees.
Common Phone: A Multilingual Dataset for Robust Acoustic Modelling (2022.lrec-1)

Copied to clipboard

Challenge: Current state-of-the-art acoustic models can easily comprise more than 100 million parameters.
Approach: They propose to train a gender-balanced, multilingual corpus from 76.000 contributors via Mozilla’s Common Voice project to perform phonetic symbol recognition and validate the quality of the generated phonetic annotation.
Outcome: The proposed model can perform phonetic symbol recognition and validate the quality of the generated phonetic annotation.
Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser Development (2024.lrec-main)

Copied to clipboard

Challenge: The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation.
Approach: They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework.
Outcome: The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators.
PRODIS - a Speech Database and a Phoneme-based Language Model for the Study of Predictability Effects in Polish (2024.lrec-main)

Copied to clipboard

Challenge: acoustic predictability is operationalised by surprisal in Polish, but cross-linguistic differences depend on prosodic system.
Approach: They present a speech database and a phoneme-level language model of Polish . they aim to study contextual predictability effects on acoustic distinctiveness .
Outcome: The proposed model is the first large, publicly available speech database of Polish . it is based on a light GPT architecture and can be expanded to other languages .
DiaBiz.Kom - towards a Polish Dialogue Act Corpus Based on ISO 24617-2 Standard (2022.coling-1)

Copied to clipboard

Challenge: DiaBiz.Kom is the first corpus of dialogue texts in the Polish language . it contains transcriptions of telephone conversations conducted according to a prepared scenario.
Approach: They describe the specification and evaluation of DiaBiz.Kom - the corpus of dialogue texts in Polish.
Outcome: The proposed corpus contains transcriptions of telephone conversations conducted according to a prepared scenario and will be used to develop a system of dialog analysis and modules for creating advanced chatbots.
Gos 2: A New Reference Corpus of Spoken Slovenian (2024.lrec-main)

Copied to clipboard

Challenge: a new corpus of spoken Slovenian has been added to the Gos reference corpus . the corpus is now more than double the original size of 300 hours, 2.4 million words .
Approach: They propose to add speech recordings and transcriptions from two related initiatives, the Gos VideoLectures corpus of public academic speech, and the Artur speech recognition database.
Outcome: The new corpus is double the original size and contains 2.4 million words . it includes speech recordings and transcriptions from two related initiatives .
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible.
Approach: They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries.
Outcome: The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations