Large Vocabulary Read Speech Corpora for Four Ethiopian Languages: Amharic, Tigrigna, Oromo and Wolaytta (2020.lrec-1)
Copied to clipboard
Solomon Teferra Abate, Martha Yifiru Tachbelie, Michael Melese, Hafte Abera, Tewodros Abebe, Wondwossen Mulugeta, Yaregal Assabie, Million Meshesha, Solomon Afnafu, Binyam Ephrem Seyoum
| Challenge: | Automatic Speech Recognition (ASR) is one of the most important technologies to support spoken communication in modern life. |
| Approach: | They have developed four large speech corpora for four Ethiopian languages . they have word error rates of 37.65%, 31.03%, 38.02%, 33.89% for each language . |
| Outcome: | The proposed corpora achieve word error rates of 37.65%, 31.03%, 38.02%, 33.89% for Amharic, Tigrigna, Oromo and Wolaytta. |
Similar Papers
Analysis of GlobalPhone and Ethiopian Languages Speech Corpora for Multilingual ASR (2020.lrec-1)
Copied to clipboard
| Challenge: | Using global phone data, we can develop multilingual speech recognition systems in yet unsupported languages. |
| Approach: | They analyze phonetic overlaps between GlobalPhone and Ethiopian speech corpora to develop multilingual Automatic Speech Recognition system for the Ethiopian languages. |
| Outcome: | The proposed system will be able to support three different languages and have morphological complexity. |
EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation (2024.lrec-main)
Copied to clipboard
Atnafu Lambebo Tonja, Israel Abebe Azime, Tadesse Destaw Belay, Mesay Gemeda Yigezu, Moges Ahmed Ah Mehamed, Abinew Ali Ayele, Ebrahim Chekol Jibril, Michael Melese Woldeyohannis, Olga Kolesnikova, Philipp Slusallek, Dietrich Klakow, Seid Muhie Yimam
| Challenge: | Low-resource languages are lagging behind current state-of-the-art (SOTA) developments in the field of NLP due to insufficient resources to train LLMs. |
| Approach: | They propose to use multilingual large language models for five Ethiopian languages and a benchmark dataset to evaluate their performance. |
| Outcome: | The proposed models outperform existing models in five Ethiopian languages and a benchmark dataset for various downstream NLP tasks. |
Parallel Corpora for bi-lingual English-Ethiopian Languages Statistical Machine Translation (C18-1)
Copied to clipboard
Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha, Solomon Atinafu, Wondwossen Mulugeta, Yaregal Assabie, Hafte Abera, Binyam Ephrem, Tewodros Abebe, Wondimagegnhue Tsegaye, Amanuel Lemma, Tsegaye Andargie, Seifedin Shifaw
| Challenge: | Various approaches to machine translation have been and are being used in the research community, that can broadly classified as rule-based and corpus based. |
| Approach: | They propose to develop parallel corpora for English and Ethiopian languages such as Amharic, Tigrigna, Afan-Oromo, Wolaytta and Ge’ez. |
| Outcome: | The proposed system improves on the English-Ethiopian languages. |
Open ASR for Icelandic: Resources and a Baseline System (L18-1)
Copied to clipboard
| Challenge: | Existing language resources are not sufficient for less-resourced languages, but a system with sufficient resources is needed. |
| Approach: | They describe available language resources and their preparation for use in a large vocabulary speech recognition system for Icelandic. |
| Outcome: | The proposed system improves on acoustic training sets and a speech corpus with a pronunciation dictionary. |
BembaSpeech: A Speech Recognition Corpus for the Bemba Language (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing speech recognition systems for African languages are very low . lack of resources (speech and text) can be attributed to poor quality of speech. |
| Approach: | They present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting of 24 hours of read speech in the Bemba language. |
| Outcome: | The proposed model achieves a word error rate (WER) of 32.91% on the Bemba language . the 1 billion XLS-R parameter model achieve better performance than the monolingual pre-trained English model on the corpus. |
An (unhelpful) guide to selecting the best ASR architecture for your under-resourced language (2023.acl-short)
Copied to clipboard
| Challenge: | English ASR now has word error rates comparable to that of human transcriptionists, but only for the handful of the world's 7000 languages with abundant training resources. |
| Approach: | They propose to use four of the most popular ASR toolkits to train ASR models for eleven languages with limited ASR training resources: eleven widely spoken languages of Africa, Asia, and South America, one endangered language of Central America, and three critically endangered languages of North America. |
| Outcome: | The proposed architecture outperforms four of the most popular ASR toolkits for eleven languages with limited training resources. |
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)
Copied to clipboard
Roberts Dargis, Arturs Znotins, Ilze Auzina, Baiba Saulite, Sanita Reinsone, Raivis Dejus, Antra Klavinska, Normunds Gruzitis
| Challenge: | Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched . |
| Approach: | a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples . |
| Outcome: | a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year. |
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in Automatic Speech Recognition (ASR) have been fueled by massive speech corpora, but extending coverage to diverse languages with limited resources remains a formidable challenge. |
| Approach: | They propose a pipeline that converts large-scale text corpora into synthetic speech using off-the-shelf text-to-speech (TTS) models. |
| Outcome: | The proposed pipeline generates 500,000 hours of synthetic speech in ten languages and achieves transcription error reductions of over 30%. |
Automatic Speech Recognition in Sanskrit: A New Speech Corpus and Modelling Insights (2021.findings-acl)
Copied to clipboard
| Challenge: | In this paper, we propose the first large scale study of automatic speech recognition in Sanskrit . we focus on the impact of unit selection in San's ASR systems . |
| Approach: | They propose a large scale study of automatic speech recognition in Sanskrit . they propose syllable level unit selection that captures character sequences . |
| Outcome: | The proposed model captures character sequences from one vowel in the word to the next vowela. |
The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing work in the area of radio browsing using automatic speech recognition (ASR) has been done by the United Nations in Uganda, and Keyword Spotting systems in Somalia. |
| Approach: | They propose to use a Luganda radio speech corpus of 155 hours to build a usable radio monitoring automatic speech recognition system. |
| Outcome: | The makerere artificial intelligence lab releases a Luganda radio speech corpus of 155 hours. |