Challenge: Existing speech FMs fall short of full compliance with open-source principles . existing models do not have model weights, code, and training data publicly available .
Approach: They propose to use a CC-BY license to create open-source speech FMs for EU languages . they collect suitable training data by surveying automatic speech recognition datasets .
Outcome: The proposed model can be used in the 24 official languages of the European Union.

Similar Papers

Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for crowdsourcing data collection require a human workforce, which is hard to sustain.
Approach: They propose to use Speech Foundation Models to automate validation processes . they find that SFMs can reduce reliance on human validation .
Outcome: The proposed model reduces the reliance on human validation without degrading the quality of the final data.
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)

Copied to clipboard

Challenge: Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched .
Approach: a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples .
Outcome: a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year.
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)

Copied to clipboard

Challenge: The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date.
Approach: They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS .
Outcome: The proposed model can build automatic speech recognition models for 700 languages.
On the Evaluation of Speech Foundation Models for Spoken Language Understanding (2024.findings-acl)

Copied to clipboard

Challenge: Spoken language understanding evaluation (SLUE) benchmarks are used to benchmark complex spoken language understanding tasks on natural speech.
Approach: They propose a set of benchmark tasks to evaluate spoken language understanding on natural speech . they use pre-trained speech foundation models to evaluate the utility of different SFMs .
Outcome: The proposed framework outperforms pre-trained speech foundation models on natural speech . the proposed framework also outperformed self-supervised SFMs on the sequence generation tasks .
“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved almost human-like performance on various tasks.
Approach: They are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate and release open-source LLMs tailored for Romanian.
Outcome: The proposed model trains, evaluates and releases open-source models tailored for Romanian.
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing training datasets for large language models are often not fully disclosed.
Approach: They propose a multilingual dataset with 6.3 trillion tokens in 167 languages . they use a pipeline of multiple stages to achieve the best quality for model training .
Outcome: The proposed dataset is cleaned and deduplicated to achieve the best quality for model training . lack of transparency has hindered research on attributing and addressing hallucination and bias issues . 6.3 trillion tokens in 167 languages are used to train multilingual LLMs .
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself.
Approach: They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations.
Outcome: The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform.
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning (2025.acl-long)

Copied to clipboard

Challenge: Recent years have witnessed significant advancements in integrating speech and audio capabilities into large language models.
Approach: They propose a 50M-example dataset for instruction fine-tuning and pre-training of speech-text large language models (LLMs) the dataset spans five languages and enables a diverse range of speech understanding and controllable speech generation instructions.
Outcome: The proposed dataset outperforms existing speech-text LLMs on instruction-following benchmarks while achieving competitive performance on foundational speech tasks.
How to Adapt Your Pretrained Multilingual Model to 1600 Languages (2021.acl-long)

Copied to clipboard

Challenge: Pretrained multilingual models perform best for languages seen during pretraining . methods exist to improve performance for unseen languages, but have been evaluated using amounts of raw text only available for a small fraction of the world’s languages.
Approach: They evaluate the performance of existing methods to adapt pretrained multilingual models to new languages using a resource available for close to 1600 languages: the New Testament.
Outcome: The proposed models perform best for languages seen during pretraining . the results show that the most efficient approach is simplest and the most accurate .
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Automatic Speech Recognition (ASR) have been fueled by massive speech corpora, but extending coverage to diverse languages with limited resources remains a formidable challenge.
Approach: They propose a pipeline that converts large-scale text corpora into synthetic speech using off-the-shelf text-to-speech (TTS) models.
Outcome: The proposed pipeline generates 500,000 hours of synthetic speech in ten languages and achieves transcription error reductions of over 30%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations