Papers by Askars Salimbajevs
Creating Lithuanian and Latvian Speech Corpora from Inaccurately Annotated Web Data (L18-1)
Copied to clipboard
| Challenge: | Existing acoustic model training data for low resource languages is not enough for low-resource languages such as Lithuanian and Latvian. |
| Approach: | They propose a method to align audio data from the Web with imprecise non-normalised transcripts for acoustic models. |
| Outcome: | The proposed method significantly improves word error rate for Lithuanian from 40% to 23% and word error rates for Latvian from 19% to 17%. |
Code-Mixed Text Augmentation for Latvian ASR (2024.lrec-main)
Copied to clipboard
| Challenge: | a new study attempts to tackle code-mixed speech recognition by improving the language model of a hybrid system. |
| Approach: | They propose an inflected transliteration and phonetic transcription model for code-mixed Latvian sentences . they leverage a large human-translated English-Latvian parallel text corpus to generate synthetic Latvian phrases . |
| Outcome: | The proposed system improves on a human-translated English-Latvian parallel text corpus . the results show that the proposed system can generate code-mixed Latvian sentences . |