Papers by Rustem Yeshpanov

5 papers
KazQAD: Kazakh Open-Domain Question Answering Dataset (2024.lrec-main)

Copied to clipboard

Challenge: KazQAD contains just under 6,000 unique questions with extracted short answers and nearly 12,000 passage-level relevance judgements.
Approach: They introduce a Kazakh open-domain question answering dataset that can be used in reading comprehension and full ODQA settings.
Outcome: The proposed dataset can be used in reading comprehension and full ODQA settings, as well as for information retrieval experiments.
KazNERD: Kazakh Named Entity Recognition Dataset (2022.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a subtask of information extraction aimed at identifying named entities (NEs) in semi-or unstructured text and classifying them into pre-specified types.
Approach: They present a dataset for Kazakh named entity recognition using an annotation scheme and guidelines for annotation.
Outcome: The dataset contains 112,702 sentences and 136,333 annotations for 25 entity classes.
KazParC: Kazakh Parallel Corpus for Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Statistical machine translation gained ground over rule-based machine translation in the late 1990s thanks to its ability to learn from large bilingual corpora.
Approach: They propose to develop a parallel corpus for machine translation across Kazakh, English, Russian, and Turkish.
Outcome: The proposed model outperforms Google Translate and Yandex Translate in terms of performance and evaluation metrics.
KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and Attitudes (2024.lrec-main)

Copied to clipboard

Challenge: Currently, sentiment analysis is a widely employed text classification task that involves extracting the sentiment expressed by individuals towards a variety of entities.
Approach: They propose to use KazSAnDRA to automate Kazakh sentiment analysis by developing and evaluating four machine learning models for polarity and score classification.
Outcome: The proposed dataset is the first and largest publicly available dataset of its kind.
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis (2024.lrec-main)

Copied to clipboard

Challenge: Using KazEmoTTS, synthesized speech still faces significant difficulties in expressing paralinguistic features such as emotions.
Approach: They created a KazEmoTTS dataset with 54,760 audio-text pairs and a TTS model trained on the KazEmpoTTs dataset.
Outcome: The proposed dataset yields an MCD score of 6.02 to 7.67 and a MOS of 3.51 to 3.57.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations