Becoming a High-Resource Language in Speech: The Catalan Case in the Common Voice Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | a project to create a publicly available voice dataset for speech recognition systems in Catalan is a multifaceted challenge. |
| Approach: | They propose to create a publicly available voice dataset for future speech technologies in Catalan using the Mozilla Common Voice crowd-sourcing platform. |
| Outcome: | The proposed dataset shows that Catalan ranks as the most prominent language in the corpus. |
Similar Papers
Common Voice: A Massively-Multilingual Speech Corpus (2020.lrec-1)
Copied to clipboard
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, Gregor Weber
| Challenge: | Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development. |
| Approach: | They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments. |
| Outcome: | The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages. |
Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? A Comprehensive Assessment for Catalan (2021.findings-acl)
Copied to clipboard
Jordi Armengol-Estapé, Casimiro Pio Carrino, Carlos Rodriguez-Penagos, Ona de Gibert Bonet, Carme Armentano-Oller, Aitor Gonzalez-Agirre, Maite Melero, Marta Villegas
| Challenge: | Multilingual language models have been a crucial breakthrough for under-resourced languages . however, the superiority of language-specific models has already been proven for underresourced ones . |
| Approach: | They propose to build a monolingual monolingual model that is comparable to state-of-the-art large multilingual models. |
| Outcome: | The proposed model consistently outperforms state-of-the-art models across tasks and settings. |
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)
Copied to clipboard
Roberts Dargis, Arturs Znotins, Ilze Auzina, Baiba Saulite, Sanita Reinsone, Raivis Dejus, Antra Klavinska, Normunds Gruzitis
| Challenge: | Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched . |
| Approach: | a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples . |
| Outcome: | a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year. |
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)
Copied to clipboard
| Challenge: | Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself. |
| Approach: | They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations. |
| Outcome: | The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform. |
Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech (2020.lrec-1)
Copied to clipboard
Adriana Guevara-Rukoz, Isin Demirsahin, Fei He, Shan-Hui Cathy Chu, Supheakmungkol Sarin, Knot Pipatsrisawat, Alexander Gutkin, Alena Butryna, Oddur Kjartansson
| Challenge: | Using crowd-sourced datasets, we build a text-to-speech voice for a new dialect in a language with existing resources. |
| Approach: | They propose a multidialectal corpus approach for building a text-to-speech voice for a new dialect in a language with existing resources using crowd-sourcing. |
| Outcome: | The proposed model outperforms baseline models in a “zero-resource” dialect scenario while holding out target dialect recordings from the training data. |
Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan (2024.lrec-main)
Copied to clipboard
Aitor Gonzalez-Agirre, Montserrat Marimon, Carlos Rodriguez-Penagos, Javier Aula-Blasco, Irene Baucells, Carme Armentano-Oller, Jorge Palomar-Giner, Baybars Kulebi, Marta Villegas
| Challenge: | Aina Project aims to provide Catalan with the resources needed to keep its relevance in AI/NLP applications. |
| Approach: | They propose a set of strategies to consider when improving technology support for a mid- or low-resource language . they propose annotated datasets and a framework to make models ready to use . |
| Outcome: | The Aina Project aims to provide Catalan with the necessary resources to keep its relevance in AI/NLP-related industry and research. |
Common Phone: A Multilingual Dataset for Robust Acoustic Modelling (2022.lrec-1)
Copied to clipboard
| Challenge: | Current state-of-the-art acoustic models can easily comprise more than 100 million parameters. |
| Approach: | They propose to train a gender-balanced, multilingual corpus from 76.000 contributors via Mozilla’s Common Voice project to perform phonetic symbol recognition and validate the quality of the generated phonetic annotation. |
| Outcome: | The proposed model can perform phonetic symbol recognition and validate the quality of the generated phonetic annotation. |
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning (2025.acl-long)
Copied to clipboard
| Challenge: | Despite their growing importance, the quality of these datasets remains under-researched. |
| Approach: | They propose guidelines and recommendations to address quality issues in future dataset development . they find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages . |
| Outcome: | The results highlight the need for proactive language planning and enhanced data quality control in the process of automatic speech recognition dataset creation. |
Samrómur: Crowd-sourcing Data Collection for Icelandic Speech Recognition (2020.lrec-1)
Copied to clipboard
David Erik Mollberg, Ólafur Helgi Jónsson, Sunneva Þorsteinsdóttir, Steinþór Steingrímsson, Eydís Huld Magnúsdóttir, Jon Gudnason
| Challenge: | Samrómur is a web application built upon Mozilla Foundation’s open-source voice collection platform Common Voice. |
| Approach: | They describe an ongoing project of speech data collection using the web application Samrómur which is built upon Common Voice, Mozilla Foundation’s web platform for open-source voice collection. |
| Outcome: | The proposed system will be the largest open speech corpus for Icelandic collected from the public domain. |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |