MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages (2024.emnlp-main)
Copied to clipboard
Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih, Matteo Negri
| Challenge: | Existing speech FMs fall short of full compliance with open-source principles . existing models do not have model weights, code, and training data publicly available . |
| Approach: | They propose to use a CC-BY license to create open-source speech FMs for EU languages . they collect suitable training data by surveying automatic speech recognition datasets . |
| Outcome: | The proposed model can be used in the 24 official languages of the European Union. |
Similar Papers
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for crowdsourcing data collection require a human workforce, which is hard to sustain. |
| Approach: | They propose to use Speech Foundation Models to automate validation processes . they find that SFMs can reduce reliance on human validation . |
| Outcome: | The proposed model reduces the reliance on human validation without degrading the quality of the final data. |
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)
Copied to clipboard
Roberts Dargis, Arturs Znotins, Ilze Auzina, Baiba Saulite, Sanita Reinsone, Raivis Dejus, Antra Klavinska, Normunds Gruzitis
| Challenge: | Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched . |
| Approach: | a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples . |
| Outcome: | a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year. |
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)
Copied to clipboard
| Challenge: | The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date. |
| Approach: | They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS . |
| Outcome: | The proposed model can build automatic speech recognition models for 700 languages. |
On the Evaluation of Speech Foundation Models for Spoken Language Understanding (2024.findings-acl)
Copied to clipboard
Siddhant Arora, Ankita Pasad, Chung-Ming Chien, Jionghao Han, Roshan Sharma, Jee-weon Jung, Hira Dhamyal, William Chen, Suwon Shon, Hung-yi Lee, Karen Livescu, Shinji Watanabe
| Challenge: | Spoken language understanding evaluation (SLUE) benchmarks are used to benchmark complex spoken language understanding tasks on natural speech. |
| Approach: | They propose a set of benchmark tasks to evaluate spoken language understanding on natural speech . they use pre-trained speech foundation models to evaluate the utility of different SFMs . |
| Outcome: | The proposed framework outperforms pre-trained speech foundation models on natural speech . the proposed framework also outperformed self-supervised SFMs on the sequence generation tasks . |
“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions (2024.findings-emnlp)
Copied to clipboard
Mihai Masala, Denis Ilie-Ablachim, Alexandru Dima, Dragos Georgian Corlatescu, Miruna-Andreea Zavelca, Ovio Olaru, Simina-Maria Terian, Andrei Terian, Marius Leordeanu, Horia Velicu, Marius Popescu, Mihai Dascalu, Traian Rebedea
| Challenge: | Large Language Models (LLMs) have achieved almost human-like performance on various tasks. |
| Approach: | They are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate and release open-source LLMs tailored for Romanian. |
| Outcome: | The proposed model trains, evaluates and releases open-source models tailored for Romanian. |
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages (2024.lrec-main)
Copied to clipboard
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, Thien Huu Nguyen
| Challenge: | Existing training datasets for large language models are often not fully disclosed. |
| Approach: | They propose a multilingual dataset with 6.3 trillion tokens in 167 languages . they use a pipeline of multiple stages to achieve the best quality for model training . |
| Outcome: | The proposed dataset is cleaned and deduplicated to achieve the best quality for model training . lack of transparency has hindered research on attributing and addressing hallucination and bias issues . 6.3 trillion tokens in 167 languages are used to train multilingual LLMs . |
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)
Copied to clipboard
| Challenge: | Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself. |
| Approach: | They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations. |
| Outcome: | The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform. |
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning (2025.acl-long)
Copied to clipboard
Prabhat Pandey, Rupak Vignesh Swaminathan, K V Vijay Girish, Arunasish Sen, Jian. Xie, Grant Strimel, Andreas Schwarz
| Challenge: | Recent years have witnessed significant advancements in integrating speech and audio capabilities into large language models. |
| Approach: | They propose a 50M-example dataset for instruction fine-tuning and pre-training of speech-text large language models (LLMs) the dataset spans five languages and enables a diverse range of speech understanding and controllable speech generation instructions. |
| Outcome: | The proposed dataset outperforms existing speech-text LLMs on instruction-following benchmarks while achieving competitive performance on foundational speech tasks. |
How to Adapt Your Pretrained Multilingual Model to 1600 Languages (2021.acl-long)
Copied to clipboard
| Challenge: | Pretrained multilingual models perform best for languages seen during pretraining . methods exist to improve performance for unseen languages, but have been evaluated using amounts of raw text only available for a small fraction of the world’s languages. |
| Approach: | They evaluate the performance of existing methods to adapt pretrained multilingual models to new languages using a resource available for close to 1600 languages: the New Testament. |
| Outcome: | The proposed models perform best for languages seen during pretraining . the results show that the most efficient approach is simplest and the most accurate . |
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in Automatic Speech Recognition (ASR) have been fueled by massive speech corpora, but extending coverage to diverse languages with limited resources remains a formidable challenge. |
| Approach: | They propose a pipeline that converts large-scale text corpora into synthetic speech using off-the-shelf text-to-speech (TTS) models. |
| Outcome: | The proposed pipeline generates 500,000 hours of synthetic speech in ten languages and achieves transcription error reductions of over 30%. |