Increasing the Accessibility of Time-Aligned Speech Corpora with Spokes Mix (L18-1)
Copied to clipboard
| Challenge: | Spokes Mix is an online service providing access to spoken corpora of Polish . high-quality corporata of conversational language are expensive to acquire . |
| Approach: | a new online service provides access to spoken corpora of Polish . the service provides a centralized, easy-to-use corpus query engine with a responsive web interface . |
| Outcome: | the proposed service provides access to spoken corpora of Polish, including three newly released time-aligned collections of manually transcribed spoken-conversational data. |
Similar Papers
An Application for Building a Polish Telephone Speech Corpus (L18-1)
Copied to clipboard
| Challenge: | Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance. |
| Approach: | They propose to build a tool for speech corpus collection of a specific domain content. |
| Outcome: | The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks. |
Processing and Understanding Mixed Language Data (D19-2)
Copied to clipboard
| Challenge: | Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text . |
| Approach: | a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities. |
| Outcome: | a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing. |
PRODIS - a Speech Database and a Phoneme-based Language Model for the Study of Predictability Effects in Polish (2024.lrec-main)
Copied to clipboard
| Challenge: | acoustic predictability is operationalised by surprisal in Polish, but cross-linguistic differences depend on prosodic system. |
| Approach: | They present a speech database and a phoneme-level language model of Polish . they aim to study contextual predictability effects on acoustic distinctiveness . |
| Outcome: | The proposed model is the first large, publicly available speech database of Polish . it is based on a light GPT architecture and can be expanded to other languages . |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |
Gos 2: A New Reference Corpus of Spoken Slovenian (2024.lrec-main)
Copied to clipboard
| Challenge: | a new corpus of spoken Slovenian has been added to the Gos reference corpus . the corpus is now more than double the original size of 300 hours, 2.4 million words . |
| Approach: | They propose to add speech recordings and transcriptions from two related initiatives, the Gos VideoLectures corpus of public academic speech, and the Artur speech recognition database. |
| Outcome: | The new corpus is double the original size and contains 2.4 million words . it includes speech recordings and transcriptions from two related initiatives . |
Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser Development (2024.lrec-main)
Copied to clipboard
Maciej Ogrodniczuk, Aleksandra Tomaszewska, Daniel Ziembicki, Sebastian Żurowski, Ryszard Tuora, Aleksandra Zwierzchowska
| Challenge: | The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation. |
| Approach: | They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework. |
| Outcome: | The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators. |
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations (2025.findings-acl)
Copied to clipboard
| Challenge: | Discourse parsing datasets based on conversations are restricted to a single domain . a lack of discourse structures in audio-based conversations is a challenge . |
| Approach: | They introduce CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse parsing in conversations. |
| Outcome: | The proposed corpus is code-mixed in Hindi and English and annotated with nine discourse relations. |
Creating a Data Set of Abstractive Summaries of Turn-labeled Spoken Human-Computer Conversations (2022.lrec-1)
Copied to clipboard
| Challenge: | Digital recorded written and spoken dialogues are becoming more available due to the growing popularity of online messenger services and chatbots. |
| Approach: | They propose to use Dutch spoken human-computer conversations, an annotation layer of turn labels, and conversational abstractive summaries of user answers to build a conversational agent. |
| Outcome: | The proposed system can be integrated into a conversational agent. |
Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech (2024.lrec-main)
Copied to clipboard
| Challenge: | The Dutch Dialect Database contains dialectal variations of Dutch recorded in the second half of the twentieth century. |
| Approach: | They propose to create a corpus containing audio recordings and orthographic transcriptions of Dutch dialects recorded in the second half of the 20th century. |
| Outcome: | The Dutch Dialect Database contains dialectal variations recorded all over the Netherlands in the second half of the twentieth century. |
MixRED: A Mix-lingual Relation Extraction Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing research focuses on monolingual relation extraction, but there is a significant gap in understanding relation extraction in the mix-lingual scenario. |
| Approach: | They propose a task of considering relation extraction in the mix-lingual scenario . they construct a human-annotated dataset to support the task . |
| Outcome: | The proposed task evaluates state-of-the-art supervised models and large language models on the human-annotated dataset MixRED. |