Papers by Rustem Yeshpanov
KazQAD: Kazakh Open-Domain Question Answering Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | KazQAD contains just under 6,000 unique questions with extracted short answers and nearly 12,000 passage-level relevance judgements. |
| Approach: | They introduce a Kazakh open-domain question answering dataset that can be used in reading comprehension and full ODQA settings. |
| Outcome: | The proposed dataset can be used in reading comprehension and full ODQA settings, as well as for information retrieval experiments. |
KazNERD: Kazakh Named Entity Recognition Dataset (2022.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a subtask of information extraction aimed at identifying named entities (NEs) in semi-or unstructured text and classifying them into pre-specified types. |
| Approach: | They present a dataset for Kazakh named entity recognition using an annotation scheme and guidelines for annotation. |
| Outcome: | The dataset contains 112,702 sentences and 136,333 annotations for 25 entity classes. |
KazParC: Kazakh Parallel Corpus for Machine Translation (2024.lrec-main)
Copied to clipboard
| Challenge: | Statistical machine translation gained ground over rule-based machine translation in the late 1990s thanks to its ability to learn from large bilingual corpora. |
| Approach: | They propose to develop a parallel corpus for machine translation across Kazakh, English, Russian, and Turkish. |
| Outcome: | The proposed model outperforms Google Translate and Yandex Translate in terms of performance and evaluation metrics. |
KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and Attitudes (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, sentiment analysis is a widely employed text classification task that involves extracting the sentiment expressed by individuals towards a variety of entities. |
| Approach: | They propose to use KazSAnDRA to automate Kazakh sentiment analysis by developing and evaluating four machine learning models for polarity and score classification. |
| Outcome: | The proposed dataset is the first and largest publicly available dataset of its kind. |
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis (2024.lrec-main)
Copied to clipboard
| Challenge: | Using KazEmoTTS, synthesized speech still faces significant difficulties in expressing paralinguistic features such as emotions. |
| Approach: | They created a KazEmoTTS dataset with 54,760 audio-text pairs and a TTS model trained on the KazEmpoTTs dataset. |
| Outcome: | The proposed dataset yields an MCD score of 6.02 to 7.67 and a MOS of 3.51 to 3.57. |