| Challenge: | Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself. |
| Approach: | They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations. |
| Outcome: | The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform. |
Similar Papers
Samrómur: Crowd-sourcing Data Collection for Icelandic Speech Recognition (2020.lrec-1)
Copied to clipboard
David Erik Mollberg, Ólafur Helgi Jónsson, Sunneva Þorsteinsdóttir, Steinþór Steingrímsson, Eydís Huld Magnúsdóttir, Jon Gudnason
| Challenge: | Samrómur is a web application built upon Mozilla Foundation’s open-source voice collection platform Common Voice. |
| Approach: | They describe an ongoing project of speech data collection using the web application Samrómur which is built upon Common Voice, Mozilla Foundation’s web platform for open-source voice collection. |
| Outcome: | The proposed system will be the largest open speech corpus for Icelandic collected from the public domain. |
SamróMur MilljóN: An ASR Corpus of One Million Verified Read Prompts in Icelandic (2024.lrec-main)
Copied to clipboard
| Challenge: | samrómur is a crowdsourcing web application designed to collect speech data for the advancement of language technologies in Icelandic. |
| Approach: | They propose to use a crowdsourcing web application to collect and verify Icelandic speech data for automatic speech recognition (ASR) they introduce a dataset comprising one million audio clips from the application . |
| Outcome: | The proposed system can produce high-quality speech data for Icelandic . the proposed system is based on a crowdsourced web application built on Mozilla's Common Voice . |
Common Voice: A Massively-Multilingual Speech Corpus (2020.lrec-1)
Copied to clipboard
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, Gregor Weber
| Challenge: | Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development. |
| Approach: | They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments. |
| Outcome: | The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages. |
Becoming a High-Resource Language in Speech: The Catalan Case in the Common Voice Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | a project to create a publicly available voice dataset for speech recognition systems in Catalan is a multifaceted challenge. |
| Approach: | They propose to create a publicly available voice dataset for future speech technologies in Catalan using the Mozilla Common Voice crowd-sourcing platform. |
| Outcome: | The proposed dataset shows that Catalan ranks as the most prominent language in the corpus. |
Don’t Rule Out Monolingual Speakers: A Method For Crowdsourcing Machine Translation Data (2021.acl-short)
Copied to clipboard
| Challenge: | High-performing machine translation systems require large amounts of training data in the form of parallel sentences, and translators are difficult to find and expensive. |
| Approach: | They propose a data collection strategy which uses graphics interchange formats (GIFs) as a pivot to collect parallel sentences from monolingual annotators. |
| Outcome: | The proposed method collects parallel sentences from monolingual annotators in Hindi, Tamil and English. |
Crowdsourced Multimodal Corpora Collection Tool (L18-1)
Copied to clipboard
| Challenge: | a crowd-sourced corpora recording method has several disadvantages, including the cost of staff, equipment and time spent recording in-lab. |
| Approach: | They propose to use a crowd-sourced data collection tool to gather controlled multimodal data of people in a rapid and scalable fashion. |
| Outcome: | The proposed tool will allow researchers to quickly gather large amounts of multimodal data spanning a wide demographic range and create their own multimodal corpus. |
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for crowdsourcing data collection require a human workforce, which is hard to sustain. |
| Approach: | They propose to use Speech Foundation Models to automate validation processes . they find that SFMs can reduce reliance on human validation . |
| Outcome: | The proposed model reduces the reliance on human validation without degrading the quality of the final data. |
Crowd-sourcing annotation of complex NLU tasks: A case study of argumentative content annotation (D19-59)
Copied to clipboard
| Challenge: | Recent advances in machine reading and listening comprehension involve the annotation of long texts. |
| Approach: | They propose a way to perform a sentence-by-sentence annotation task with crowd annotators. |
| Outcome: | The proposed approach can be used to identify claims in a debate speech. |
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)
Copied to clipboard
Roberts Dargis, Arturs Znotins, Ilze Auzina, Baiba Saulite, Sanita Reinsone, Raivis Dejus, Antra Klavinska, Normunds Gruzitis
| Challenge: | Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched . |
| Approach: | a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples . |
| Outcome: | a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year. |
Samrómur Children: An Icelandic Speech Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Samrómur Children contains 131 hours of read speech from Icelandic children aged between 4 to 17 years. |
| Approach: | They propose to build a large-scale speech corpus for automatic speech recognition for Icelandic. |
| Outcome: | The corpus contains 131 hours of read speech from Icelandic children aged 4 to 17 years . the goal of the project is to make Icelandic available in language-technology applications . |