Challenge: samrómur is a crowdsourcing web application designed to collect speech data for the advancement of language technologies in Icelandic.
Approach: They propose to use a crowdsourcing web application to collect and verify Icelandic speech data for automatic speech recognition (ASR) they introduce a dataset comprising one million audio clips from the application .
Outcome: The proposed system can produce high-quality speech data for Icelandic . the proposed system is based on a crowdsourced web application built on Mozilla's Common Voice .

Similar Papers

Samrómur: Crowd-sourcing Data Collection for Icelandic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Samrómur is a web application built upon Mozilla Foundation’s open-source voice collection platform Common Voice.
Approach: They describe an ongoing project of speech data collection using the web application Samrómur which is built upon Common Voice, Mozilla Foundation’s web platform for open-source voice collection.
Outcome: The proposed system will be the largest open speech corpus for Icelandic collected from the public domain.
Samrómur Children: An Icelandic Speech Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur Children contains 131 hours of read speech from Icelandic children aged between 4 to 17 years.
Approach: They propose to build a large-scale speech corpus for automatic speech recognition for Icelandic.
Outcome: The corpus contains 131 hours of read speech from Icelandic children aged 4 to 17 years . the goal of the project is to make Icelandic available in language-technology applications .
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself.
Approach: They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations.
Outcome: The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform.
Open ASR for Icelandic: Resources and a Baseline System (L18-1)

Copied to clipboard

Challenge: Existing language resources are not sufficient for less-resourced languages, but a system with sufficient resources is needed.
Approach: They describe available language resources and their preparation for use in a large vocabulary speech recognition system for Icelandic.
Outcome: The proposed system improves on acoustic training sets and a speech corpus with a pronunciation dictionary.
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)

Copied to clipboard

Challenge: Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched .
Approach: a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples .
Outcome: a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year.
Development of Automatic Speech Recognition for the Documentation of Cook Islands Māori (2022.lrec-1)

Copied to clipboard

Challenge: a new study describes the process of data processing and training of an automatic speech recognition system for Cook Islands Mori . the system is based on statistical and Deep Learning techniques, and is available under a license .
Approach: They describe the process of data processing and training of an automatic speech recognition system for Cook Islands Mori . they transcribed four hours of speech from adults and elderly speakers of the language and prepared two experiments .
Outcome: The proposed system can perform better with low-resource Indigenous languages . the system can be used to accelerate the documentation of Cook Islands Mori .
MASRI-HEADSET: A Maltese Corpus for Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Maltese is the national language of Malta and is spoken by approximately 500,000 people.
Approach: They present the first spoken Maltese corpus designed purposely for Automatic Speech Recognition (ASR) it consists of 8 hours of speech paired with text, recorded by using short text snippets in a laboratory environment.
Outcome: The MASRI-HEADSET corpus was developed by the MASR project at the University of Malta.
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement (2025.acl-long)

Copied to clipboard

Challenge: GigaSpeech 2 is a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages.
Approach: They propose a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages and an automated pipeline for data crawling, transcription, and label refinement.
Outcome: The proposed corpus reduces the word error rate for Thai, Indonesian, and Vietnamese on a realistic YouTube test set by 25% to 40% compared to Whisper large-v3.
ÌròyìnSpeech: A Multi-purpose Yorùbá Speech Corpus (2024.lrec-main)

Copied to clipboard

Challenge: rynSpeech corpus is a dataset that can be used for both Text-to-Speecher (TTS) and Automatic Speech Recognition (ASR) speakers of many African languages have no access to voice-enabled applications in their native languages.
Approach: They propose a dataset to collect Yorùbá speech data that can be used for both TTS and ASR tasks.
Outcome: The proposed dataset can generate a good quality model with as little as 5 hours of speech . the results are consistent with previous studies on the Yorùbá language .
Speak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio Recordings (2022.lrec-1)

Copied to clipboard

Challenge: Speak is a toolkit that allows researchers to crowdsource speech recordings using Amazon Mechanical Turk (MTurk).
Approach: They propose to use Amazon Mechanical Turk to crowdsource speech recordings . they use various measures to ensure that the recordings are of adequate quality .
Outcome: Speak is an open-source toolkit that allows researchers to crowdsource speech recordings using Amazon Mechanical Turk (MTurk).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations