Development of Automatic Speech Recognition for the Documentation of Cook Islands Māori (2022.lrec-1)
Copied to clipboard
Rolando Coto-Solano, Sally Akevai Nicholas, Samiha Datta, Victoria Quint, Piripi Wills, Emma Ngakuravaru Powell, Liam Koka’ua, Syed Tanveer, Isaac Feldman
| Challenge: | a new study describes the process of data processing and training of an automatic speech recognition system for Cook Islands Mori . the system is based on statistical and Deep Learning techniques, and is available under a license . |
| Approach: | They describe the process of data processing and training of an automatic speech recognition system for Cook Islands Mori . they transcribed four hours of speech from adults and elderly speakers of the language and prepared two experiments . |
| Outcome: | The proposed system can perform better with low-resource Indigenous languages . the system can be used to accelerate the documentation of Cook Islands Mori . |
Similar Papers
Development of Community-Oriented Text-to-Speech Models for Māori ‘Avaiki Nui (Cook Islands Māori) (2024.lrec-main)
Copied to clipboard
Jesin James, Rolando Coto-Solano, Sally Akevai Nicholas, Joshua Zhu, Bovey Yu, Fuki Babasaki, Jenny Tyler Wang, Nicholas Derby
| Challenge: | Text-to-speech synthesis is used to transform text into a synthesized voice for a specific language. |
| Approach: | They describe the development of a text-to-speech system for Mori ‘Avaiki Nui (Cook Islands Mi) they used two approaches to train the system, the HMM-system MaryTTS and the deep learning system FastSpeech2 . |
| Outcome: | The proposed system is based on the HMM-system MaryTTS and the deep learning system FastSpeech2 . the ground truth voice had the highest quality, but the fastspeech 2 voice had a significantly higher quality than the MaryTTs synthesized recordings. |
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets for automatic speech recognition (ASR) in the endangered Kichwa language have been limited. |
| Approach: | They present Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador. |
| Outcome: | The proposed dataset shows that it can be used to build an automatic speech recognition system for the endangered language with reliable quality despite its small size. |
Fine-Tuning a Pre-Trained Wav2Vec2 Model for Automatic Speech Recognition- Experiments with De Zahrar Sproche (2024.lrec-main)
Copied to clipboard
| Challenge: | Developing semi-automatic methods of transcription and annotation based on small amounts of annotated data would free field linguists to focus on tasks that are linguistically and relationally significant during fieldwork. |
| Approach: | They propose to use a pre-trained model to tune a generic pre-trainer model to reduce the transcription workload of field linguists. |
| Outcome: | The proposed system reduces the transcription workload of field linguists by averaging a pre-trained model with a language-specific tuning. |
BembaSpeech: A Speech Recognition Corpus for the Bemba Language (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing speech recognition systems for African languages are very low . lack of resources (speech and text) can be attributed to poor quality of speech. |
| Approach: | They present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting of 24 hours of read speech in the Bemba language. |
| Outcome: | The proposed model achieves a word error rate (WER) of 32.91% on the Bemba language . the 1 billion XLS-R parameter model achieve better performance than the monolingual pre-trained English model on the corpus. |
Discourse on ASR Measurement: Introducing the ARPOCA Assessment Tool (2022.acl-srw)
Copied to clipboard
| Challenge: | Automated speech recognition (ASR) models are based on a corpus of audio recordings, but are often small or nonexistent for less common languages and dialects. |
| Approach: | This research proposal will develop a semi-automatic acoustic features extraction system that integrates phonetic transcripts with pronunciation dictionaries. |
| Outcome: | The proposed system will be used to improve language recognition and model feedback in less common languages and dialects. |
Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation (2023.acl-long)
Copied to clipboard
| Challenge: | Using self-training or text-to-speech (TTS) to improve low-resource ASR performance is costly and can lead to catastrophic forgetting. |
| Approach: | They examine whether data augmentation techniques could help improve low-resource ASR performance . they use self-training to generate transcriptions, which are combined with original data to train new system . |
| Outcome: | The proposed approach yields a 20.5% reduction in WER compared to a system trained on 24 minutes of manually transcribed speech. |
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)
Copied to clipboard
| Challenge: | Using acoustic data, we develop automatic speech recognition systems for three low resource languages. |
| Approach: | They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework. |
| Outcome: | The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut. |
Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset (2022.lrec-1)
Copied to clipboard
Tiezheng Yu, Rita Frieske, Peng Xu, Samuel Cahyawijaya, Cheuk Tung Yiu, Holy Lovenia, Wenliang Dai, Elham J. Barezi, Qifeng Chen, Xiaojuan Ma, Bertram Shi, Pascale Fung
| Challenge: | In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language . due to the popularization of deep learning, ASR technology has led to a significant improvement in recognizing many languages. |
| Approach: | They propose to use a dataset to analyze the data available for the Hong Kong Cantonese language . they use zh-HK as a source and a state-of-the-art ASR model to build a powerful model . |
| Outcome: | The proposed model improves on the biggest existing dataset, Common Voice zh-HK. |
Multilingual Models for ASR in Chibchan Languages (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing algorithms for low resource-intensive languages are not available for these languages . a paper comparing the performance of different models and algorithms for these extremely low resource languages is presented. |
| Approach: | They propose to fine-tune four ASR algorithms to create monolingual models for Bribri and Cabécar . they then use the best performing algorithm to train joint and transfer learning models for both languages . |
| Outcome: | The proposed algorithms are effective in both Bribri and Cabécar, but especially in Bribri. |
Fine-Tuning ASR models for Very Low-Resource Languages: A Study on Mvskoke (2024.acl-srw)
Copied to clipboard
| Challenge: | Recent advances in multilingual models for automatic speech recognition (ASR) have been able to achieve a high accuracy for languages with extremely limited resources. |
| Approach: | They examine the parameter efficiency of training an adapter for the Mvskoke language, an indigenous language of America. |
| Outcome: | The proposed model is parameter efficient and gives higher accuracy for a relatively small amount of data. |