Challenge: a project to create a publicly available voice dataset for speech recognition systems in Catalan is a multifaceted challenge.
Approach: They propose to create a publicly available voice dataset for future speech technologies in Catalan using the Mozilla Common Voice crowd-sourcing platform.
Outcome: The proposed dataset shows that Catalan ranks as the most prominent language in the corpus.

Similar Papers

Common Voice: A Massively-Multilingual Speech Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development.
Approach: They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments.
Outcome: The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages.
Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? A Comprehensive Assessment for Catalan (2021.findings-acl)

Copied to clipboard

Challenge: Multilingual language models have been a crucial breakthrough for under-resourced languages . however, the superiority of language-specific models has already been proven for underresourced ones .
Approach: They propose to build a monolingual monolingual model that is comparable to state-of-the-art large multilingual models.
Outcome: The proposed model consistently outperforms state-of-the-art models across tasks and settings.
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)

Copied to clipboard

Challenge: Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched .
Approach: a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples .
Outcome: a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year.
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself.
Approach: They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations.
Outcome: The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform.
Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech (2020.lrec-1)

Copied to clipboard

Challenge: Using crowd-sourced datasets, we build a text-to-speech voice for a new dialect in a language with existing resources.
Approach: They propose a multidialectal corpus approach for building a text-to-speech voice for a new dialect in a language with existing resources using crowd-sourcing.
Outcome: The proposed model outperforms baseline models in a “zero-resource” dialect scenario while holding out target dialect recordings from the training data.
Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan (2024.lrec-main)

Copied to clipboard

Challenge: Aina Project aims to provide Catalan with the resources needed to keep its relevance in AI/NLP applications.
Approach: They propose a set of strategies to consider when improving technology support for a mid- or low-resource language . they propose annotated datasets and a framework to make models ready to use .
Outcome: The Aina Project aims to provide Catalan with the necessary resources to keep its relevance in AI/NLP-related industry and research.
Common Phone: A Multilingual Dataset for Robust Acoustic Modelling (2022.lrec-1)

Copied to clipboard

Challenge: Current state-of-the-art acoustic models can easily comprise more than 100 million parameters.
Approach: They propose to train a gender-balanced, multilingual corpus from 76.000 contributors via Mozilla’s Common Voice project to perform phonetic symbol recognition and validate the quality of the generated phonetic annotation.
Outcome: The proposed model can perform phonetic symbol recognition and validate the quality of the generated phonetic annotation.
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning (2025.acl-long)

Copied to clipboard

Challenge: Despite their growing importance, the quality of these datasets remains under-researched.
Approach: They propose guidelines and recommendations to address quality issues in future dataset development . they find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages .
Outcome: The results highlight the need for proactive language planning and enhanced data quality control in the process of automatic speech recognition dataset creation.
Samrómur: Crowd-sourcing Data Collection for Icelandic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Samrómur is a web application built upon Mozilla Foundation’s open-source voice collection platform Common Voice.
Approach: They describe an ongoing project of speech data collection using the web application Samrómur which is built upon Common Voice, Mozilla Foundation’s web platform for open-source voice collection.
Outcome: The proposed system will be the largest open speech corpus for Icelandic collected from the public domain.
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)

Copied to clipboard

Challenge: Language is a powerful means of communication and should be regarded as more than just a collection of tokens.
Approach: They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices.
Outcome: The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations