Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself.
Approach: They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations.
Outcome: The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform.

Similar Papers

Samrómur: Crowd-sourcing Data Collection for Icelandic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Samrómur is a web application built upon Mozilla Foundation’s open-source voice collection platform Common Voice.
Approach: They describe an ongoing project of speech data collection using the web application Samrómur which is built upon Common Voice, Mozilla Foundation’s web platform for open-source voice collection.
Outcome: The proposed system will be the largest open speech corpus for Icelandic collected from the public domain.
SamróMur MilljóN: An ASR Corpus of One Million Verified Read Prompts in Icelandic (2024.lrec-main)

Copied to clipboard

Challenge: samrómur is a crowdsourcing web application designed to collect speech data for the advancement of language technologies in Icelandic.
Approach: They propose to use a crowdsourcing web application to collect and verify Icelandic speech data for automatic speech recognition (ASR) they introduce a dataset comprising one million audio clips from the application .
Outcome: The proposed system can produce high-quality speech data for Icelandic . the proposed system is based on a crowdsourced web application built on Mozilla's Common Voice .
Common Voice: A Massively-Multilingual Speech Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development.
Approach: They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments.
Outcome: The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages.
Becoming a High-Resource Language in Speech: The Catalan Case in the Common Voice Corpus (2024.lrec-main)

Copied to clipboard

Challenge: a project to create a publicly available voice dataset for speech recognition systems in Catalan is a multifaceted challenge.
Approach: They propose to create a publicly available voice dataset for future speech technologies in Catalan using the Mozilla Common Voice crowd-sourcing platform.
Outcome: The proposed dataset shows that Catalan ranks as the most prominent language in the corpus.
Don’t Rule Out Monolingual Speakers: A Method For Crowdsourcing Machine Translation Data (2021.acl-short)

Copied to clipboard

Challenge: High-performing machine translation systems require large amounts of training data in the form of parallel sentences, and translators are difficult to find and expensive.
Approach: They propose a data collection strategy which uses graphics interchange formats (GIFs) as a pivot to collect parallel sentences from monolingual annotators.
Outcome: The proposed method collects parallel sentences from monolingual annotators in Hindi, Tamil and English.
Crowdsourced Multimodal Corpora Collection Tool (L18-1)

Copied to clipboard

Challenge: a crowd-sourced corpora recording method has several disadvantages, including the cost of staff, equipment and time spent recording in-lab.
Approach: They propose to use a crowd-sourced data collection tool to gather controlled multimodal data of people in a rapid and scalable fashion.
Outcome: The proposed tool will allow researchers to quickly gather large amounts of multimodal data spanning a wide demographic range and create their own multimodal corpus.
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for crowdsourcing data collection require a human workforce, which is hard to sustain.
Approach: They propose to use Speech Foundation Models to automate validation processes . they find that SFMs can reduce reliance on human validation .
Outcome: The proposed model reduces the reliance on human validation without degrading the quality of the final data.
Crowd-sourcing annotation of complex NLU tasks: A case study of argumentative content annotation (D19-59)

Copied to clipboard

Challenge: Recent advances in machine reading and listening comprehension involve the annotation of long texts.
Approach: They propose a way to perform a sentence-by-sentence annotation task with crowd annotators.
Outcome: The proposed approach can be used to identify claims in a debate speech.
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)

Copied to clipboard

Challenge: Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched .
Approach: a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples .
Outcome: a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year.
Samrómur Children: An Icelandic Speech Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur Children contains 131 hours of read speech from Icelandic children aged between 4 to 17 years.
Approach: They propose to build a large-scale speech corpus for automatic speech recognition for Icelandic.
Outcome: The corpus contains 131 hours of read speech from Icelandic children aged 4 to 17 years . the goal of the project is to make Icelandic available in language-technology applications .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations