SamróMur MilljóN: An ASR Corpus of One Million Verified Read Prompts in Icelandic (2024.lrec-main)
Copied to clipboard
| Challenge: | samrómur is a crowdsourcing web application designed to collect speech data for the advancement of language technologies in Icelandic. |
| Approach: | They propose to use a crowdsourcing web application to collect and verify Icelandic speech data for automatic speech recognition (ASR) they introduce a dataset comprising one million audio clips from the application . |
| Outcome: | The proposed system can produce high-quality speech data for Icelandic . the proposed system is based on a crowdsourced web application built on Mozilla's Common Voice . |
Similar Papers
Samrómur: Crowd-sourcing Data Collection for Icelandic Speech Recognition (2020.lrec-1)
Copied to clipboard
David Erik Mollberg, Ólafur Helgi Jónsson, Sunneva Þorsteinsdóttir, Steinþór Steingrímsson, Eydís Huld Magnúsdóttir, Jon Gudnason
| Challenge: | Samrómur is a web application built upon Mozilla Foundation’s open-source voice collection platform Common Voice. |
| Approach: | They describe an ongoing project of speech data collection using the web application Samrómur which is built upon Common Voice, Mozilla Foundation’s web platform for open-source voice collection. |
| Outcome: | The proposed system will be the largest open speech corpus for Icelandic collected from the public domain. |
Samrómur Children: An Icelandic Speech Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Samrómur Children contains 131 hours of read speech from Icelandic children aged between 4 to 17 years. |
| Approach: | They propose to build a large-scale speech corpus for automatic speech recognition for Icelandic. |
| Outcome: | The corpus contains 131 hours of read speech from Icelandic children aged 4 to 17 years . the goal of the project is to make Icelandic available in language-technology applications . |
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)
Copied to clipboard
| Challenge: | Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself. |
| Approach: | They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations. |
| Outcome: | The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform. |
Open ASR for Icelandic: Resources and a Baseline System (L18-1)
Copied to clipboard
| Challenge: | Existing language resources are not sufficient for less-resourced languages, but a system with sufficient resources is needed. |
| Approach: | They describe available language resources and their preparation for use in a large vocabulary speech recognition system for Icelandic. |
| Outcome: | The proposed system improves on acoustic training sets and a speech corpus with a pronunciation dictionary. |
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)
Copied to clipboard
Roberts Dargis, Arturs Znotins, Ilze Auzina, Baiba Saulite, Sanita Reinsone, Raivis Dejus, Antra Klavinska, Normunds Gruzitis
| Challenge: | Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched . |
| Approach: | a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples . |
| Outcome: | a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year. |
Development of Automatic Speech Recognition for the Documentation of Cook Islands Māori (2022.lrec-1)
Copied to clipboard
Rolando Coto-Solano, Sally Akevai Nicholas, Samiha Datta, Victoria Quint, Piripi Wills, Emma Ngakuravaru Powell, Liam Koka’ua, Syed Tanveer, Isaac Feldman
| Challenge: | a new study describes the process of data processing and training of an automatic speech recognition system for Cook Islands Mori . the system is based on statistical and Deep Learning techniques, and is available under a license . |
| Approach: | They describe the process of data processing and training of an automatic speech recognition system for Cook Islands Mori . they transcribed four hours of speech from adults and elderly speakers of the language and prepared two experiments . |
| Outcome: | The proposed system can perform better with low-resource Indigenous languages . the system can be used to accelerate the documentation of Cook Islands Mori . |
MASRI-HEADSET: A Maltese Corpus for Speech Recognition (2020.lrec-1)
Copied to clipboard
Carlos Daniel Hernandez Mena, Albert Gatt, Andrea DeMarco, Claudia Borg, Lonneke van der Plas, Amanda Muscat, Ian Padovani
| Challenge: | Maltese is the national language of Malta and is spoken by approximately 500,000 people. |
| Approach: | They present the first spoken Maltese corpus designed purposely for Automatic Speech Recognition (ASR) it consists of 8 hours of speech paired with text, recorded by using short text snippets in a laboratory environment. |
| Outcome: | The MASRI-HEADSET corpus was developed by the MASR project at the University of Malta. |
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement (2025.acl-long)
Copied to clipboard
Yifan Yang, Zheshu Song, Jianheng Zhuo, Mingyu Cui, Jinpeng Li, Bo Yang, Yexing Du, Ziyang Ma, Xunying Liu, Ziyuan Wang, Ke Li, Shuai Fan, Kai Yu, Wei-Qiang Zhang, Guoguo Chen, Xie Chen
| Challenge: | GigaSpeech 2 is a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages. |
| Approach: | They propose a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages and an automated pipeline for data crawling, transcription, and label refinement. |
| Outcome: | The proposed corpus reduces the word error rate for Thai, Indonesian, and Vietnamese on a realistic YouTube test set by 25% to 40% compared to Whisper large-v3. |
ÌròyìnSpeech: A Multi-purpose Yorùbá Speech Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | rynSpeech corpus is a dataset that can be used for both Text-to-Speecher (TTS) and Automatic Speech Recognition (ASR) speakers of many African languages have no access to voice-enabled applications in their native languages. |
| Approach: | They propose a dataset to collect Yorùbá speech data that can be used for both TTS and ASR tasks. |
| Outcome: | The proposed dataset can generate a good quality model with as little as 5 hours of speech . the results are consistent with previous studies on the Yorùbá language . |
Speak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio Recordings (2022.lrec-1)
Copied to clipboard
| Challenge: | Speak is a toolkit that allows researchers to crowdsource speech recordings using Amazon Mechanical Turk (MTurk). |
| Approach: | They propose to use Amazon Mechanical Turk to crowdsource speech recordings . they use various measures to ensure that the recordings are of adequate quality . |
| Outcome: | Speak is an open-source toolkit that allows researchers to crowdsource speech recordings using Amazon Mechanical Turk (MTurk). |