Papers by Mukhamet Nurpeiissov

1 papers
A Crowdsourced Open-Source Kazakh Speech Corpus and Initial Speech Recognition Baseline (2021.eacl-main)

Copied to clipboard

Challenge: The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders.
Approach: They propose to build an open-source Kazakh speech corpus for the Kazakh language that contains over 153,000 transcribed audio . they describe the data collection and preprocessing procedures followed by a description of the database specifications.
Outcome: The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations