Papers by Askars Salimbajevs

2 papers
Creating Lithuanian and Latvian Speech Corpora from Inaccurately Annotated Web Data (L18-1)

Copied to clipboard

Challenge: Existing acoustic model training data for low resource languages is not enough for low-resource languages such as Lithuanian and Latvian.
Approach: They propose a method to align audio data from the Web with imprecise non-normalised transcripts for acoustic models.
Outcome: The proposed method significantly improves word error rate for Lithuanian from 40% to 23% and word error rates for Latvian from 19% to 17%.
Code-Mixed Text Augmentation for Latvian ASR (2024.lrec-main)

Copied to clipboard

Challenge: a new study attempts to tackle code-mixed speech recognition by improving the language model of a hybrid system.
Approach: They propose an inflected transliteration and phonetic transcription model for code-mixed Latvian sentences . they leverage a large human-translated English-Latvian parallel text corpus to generate synthetic Latvian phrases .
Outcome: The proposed system improves on a human-translated English-Latvian parallel text corpus . the results show that the proposed system can generate code-mixed Latvian sentences .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations