BembaSpeech: A Speech Recognition Corpus for the Bemba Language (2022.lrec-1)

Copied to clipboard

Challenge: Existing speech recognition systems for African languages are very low . lack of resources (speech and text) can be attributed to poor quality of speech.
Approach: They present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting of 24 hours of read speech in the Bemba language.
Outcome: The proposed model achieves a word error rate (WER) of 32.91% on the Bemba language . the 1 billion XLS-R parameter model achieve better performance than the monolingual pre-trained English model on the corpus.

Similar Papers

An Exploration of Mamba for Speech Self-Supervised Models (2026.acl-long)

Copied to clipboard

Challenge: Mamba-based SSL models are promising for long-sequence modeling, speech unit extraction, and speech self-supervised learning.
Approach: They propose to use Mamba-based HuBERT models as an alternative to Transformer-based SSL architectures.
Outcome: The proposed models outperform Transformer-based models in language modeling tasks while showing superior performance on streaming ASR.
BIG-C: a Multimodal Multi-Purpose Dataset for Bemba (2023.acl-long)

Copied to clipboard

Challenge: Bemba is the most populous language of Zambia but lacks resources for research . despite its significance, Bemba remains under-resourced and lacking in high-quality data and resources for NLP experiments and language technologies.
Approach: They propose a large multimodal dataset for Bemba that includes images, transcriptions and translations.
Outcome: The proposed dataset is based on images, transcriptions and translations of Bemba speakers . it provides baselines on speech recognition, machine translation and speech translation tasks .
The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Existing work in the area of radio browsing using automatic speech recognition (ASR) has been done by the United Nations in Uganda, and Keyword Spotting systems in Somalia.
Approach: They propose to use a Luganda radio speech corpus of 155 hours to build a usable radio monitoring automatic speech recognition system.
Outcome: The makerere artificial intelligence lab releases a Luganda radio speech corpus of 155 hours.
Large Vocabulary Read Speech Corpora for Four Ethiopian Languages: Amharic, Tigrigna, Oromo and Wolaytta (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is one of the most important technologies to support spoken communication in modern life.
Approach: They have developed four large speech corpora for four Ethiopian languages . they have word error rates of 37.65%, 31.03%, 38.02%, 33.89% for each language .
Outcome: The proposed corpora achieve word error rates of 37.65%, 31.03%, 38.02%, 33.89% for Amharic, Tigrigna, Oromo and Wolaytta.
Development of Automatic Speech Recognition for the Documentation of Cook Islands Māori (2022.lrec-1)

Copied to clipboard

Challenge: a new study describes the process of data processing and training of an automatic speech recognition system for Cook Islands Mori . the system is based on statistical and Deep Learning techniques, and is available under a license .
Approach: They describe the process of data processing and training of an automatic speech recognition system for Cook Islands Mori . they transcribed four hours of speech from adults and elderly speakers of the language and prepared two experiments .
Outcome: The proposed system can perform better with low-resource Indigenous languages . the system can be used to accelerate the documentation of Cook Islands Mori .
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation (L18-1)

Copied to clipboard

Challenge: Recent work in spoken language translation (SLT) has attempted to build end-to-end speech-totext translation without using source language transcription during learning or decoding.
Approach: They propose to augment an existing (monolingual) corpus: LibriSpeech.
Outcome: The proposed corpus is derived from read audiobooks from the LibriVox project and has been carefully segmented and aligned.
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement (2025.acl-long)

Copied to clipboard

Challenge: GigaSpeech 2 is a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages.
Approach: They propose a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages and an automated pipeline for data crawling, transcription, and label refinement.
Outcome: The proposed corpus reduces the word error rate for Thai, Indonesian, and Vietnamese on a realistic YouTube test set by 25% to 40% compared to Whisper large-v3.
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)

Copied to clipboard

Challenge: Using acoustic data, we develop automatic speech recognition systems for three low resource languages.
Approach: They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework.
Outcome: The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut.
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Automatic Speech Recognition (ASR) have been fueled by massive speech corpora, but extending coverage to diverse languages with limited resources remains a formidable challenge.
Approach: They propose a pipeline that converts large-scale text corpora into synthetic speech using off-the-shelf text-to-speech (TTS) models.
Outcome: The proposed pipeline generates 500,000 hours of synthetic speech in ten languages and achieves transcription error reductions of over 30%.
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models.
Approach: They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages.
Outcome: The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations