Papers with Kazakh

16 papers
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation (2021.eacl-srw)

Copied to clipboard

Challenge: Current NMT systems typically operate at the level of subwords, causing problems of vocabulary sparsity.
Approach: They compare subword segmentation methods with morphologically-based methods in a low-resource setting . they find that no consistent and reliable differences emerge between the methods .
Outcome: The proposed methods outperform BPE in a low-resource translation setting.
KazNERD: Kazakh Named Entity Recognition Dataset (2022.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a subtask of information extraction aimed at identifying named entities (NEs) in semi-or unstructured text and classifying them into pre-specified types.
Approach: They present a dataset for Kazakh named entity recognition using an annotation scheme and guidelines for annotation.
Outcome: The dataset contains 112,702 sentences and 136,333 annotations for 25 entity classes.
Harnessing Multilinguality in Unsupervised Machine Translation for Rare Languages (2021.naacl-main)

Copied to clipboard

Challenge: Unsupervised translation systems have impressive performance on resource-rich language pairs . however, in more realistic settings, unsupervised systems perform poorly .
Approach: They propose a model for 5 low-resource languages that leverages monolingual and auxiliary parallel data from other high-resourced languages.
Outcome: The proposed model outperforms state-of-the-art models on low-resource languages . it also matches the current state- of-the art model for Nepali-English .
Kardeş-NLU: Transfer to Low-Resource Languages with Big Brother’s Help – A Benchmark and Evaluation for Turkic Languages (2024.eacl-long)

Copied to clipboard

Challenge: Cross-lingual transfer (XLT) driven by massively multilingual language models (mmLMs) has been shown to be ineffective for low-resource (LR) target languages with little (or no) representation in mmLM’s pretraining .
Approach: They propose a benchmark to evaluate cross-lingual transfer (XLT) to LR languages that do have a close HR relative and a framework to integrate Turkish into XLT.
Outcome: The proposed configuration is of practical relevance for more of the world’s languages: XLT to LR languages that do have a close HR relative.
Uncertainty-Aware Cross-Lingual Transfer with Pseudo Partial Labels (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods to train pre-trained language models for zero-shot cross-lingual tasks are noisy and lack confidence.
Approach: They propose an uncertainty-aware cross-lingual transfer framework with pseudo-partial-label to maximize the utilization of unlabeled data by reducing noise.
Outcome: The proposed framework outperforms baselines on named entity recognition and natural language inference tasks on 40 languages.
MC2: Towards Transparent and Culturally-Aware NLP for Minority Languages in China (2024.acl-long)

Copied to clipboard

Challenge: MC2 is the largest open-source corpus of minority languages in china . MC2, however, includes four underrepresented languages: Tibetan, Uyghur, Kazakh, and Mongolian .
Approach: They propose a multilingual corpus of minority languages in China that includes four underrepresented languages . they prioritize accuracy while enhancing diversity by using a quality-centric approach .
Outcome: The proposed model prioritizes accuracy while enhancing diversity, the authors say . MC2 includes four underrepresented languages: Tibetan, Uyghur, Kazakh, and Mongolian .
Discriminating between Similar Languages on Imbalanced Conversational Texts (L18-1)

Copied to clipboard

Challenge: Empirical results suggest that our system achieves an accuracy of 95.7% on our Uyghur and Kazakh dataset, which is higher than that of the CNN classifier.
Approach: They propose to build a balanced Uyghur and Kazakh corpus and build morphological classifiers to discriminate between the two languages.
Outcome: The proposed system outperforms the champions on both test sets B1 and B2.
Cross-Lingual Word Embeddings for Turkic Languages (2020.lrec-1)

Copied to clipboard

Challenge: Existing techniques to align monolingual embeddings are difficult to use because of low resources.
Approach: They propose to use existing techniques to align monolingual embedding spaces for Turkic, Uzbek, Azeri, Kazakh and Kyrgyz languages.
Outcome: The proposed techniques outperform existing techniques on bilingual dictionaries and an extrinsic task.
Qorǵau: Evaluating Safety in Kazakh-Russian Bilingual Contexts (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have the potential to generate harmful content, posing risks to users.
Approach: They propose a dataset specifically designed for safety evaluation in Kazakh and Russian . they use a bilingual context in Kazakhstan where both Kazakh (a low-resource language) and Russian (a high-resourced language)
Outcome: The proposed dataset is designed for safety evaluation in Kazakh and Russian . it shows that both multilingual and language-specific LLMs perform better than others .
KazakhTTS2: Extending the Open-Source Kazakh TTS Corpus With More Data, Speakers, and Topics (2022.lrec-1)

Copied to clipboard

Challenge: Text-to-speech (TTS) is a process of converting written text into speech.
Approach: They present an expanded version of their text-to-speech corpus for Kazakh . they propose to use the corpus to build high-quality TTS systems for the language .
Outcome: The constructed corpus is sufficient to build robust TTS models for Kazakh and other Turkic languages, with a subjective mean opinion score ranging from 3.6 to 4.2 for all the five speakers.
MiLiC-Eval: Benchmarking Multilingual LLMs for China’s Minority Languages (2025.findings-acl)

Copied to clipboard

Challenge: Large language models excel in high-resource languages but struggle with low-resourced languages . minority languages such as Tibetan, Uyghur, Kazakh, and Mongolian are marginalized in NLP research due to limited digital representation and the scarcity of training data.
Approach: They propose a benchmark for minority languages in China that tracks the progress of large language models on low-resource languages.
Outcome: The proposed benchmark focuses on underrepresented writing systems and syntax-intensive tasks.
Stereotype Bias in a Bilingual Setting: A Culturally Grounded Evaluation in Kazakhstan (2026.acl-long)

Copied to clipboard

Challenge: Stereotype bias in language models is largely understudied in English . language models perform strongly on downstream NLP tasks, but they are pre-trained on large text corpora .
Approach: They use a dataset to assess stereotype bias in language models in Kazakhstan . they find that stereotype bias is most pronounced in code-mixed inputs .
Outcome: The proposed dataset shows that stereotype bias is most pronounced in code-mixed inputs.
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan (2025.acl-long)

Copied to clipboard

Challenge: Kazakh language remains underrepresented in the field of natural language processing despite the country's population exceeding twenty million . however, there is a lack of dedicated models and benchmark evaluations specifically tailored to Kazakh languages.
Approach: They propose to create a dataset specifically designed for Kazakh language with 23,000 questions sourced from authentic educational materials and manually validated by native speakers and educators.
Outcome: The first MMLU-style dataset specifically designed for Kazakh language.
KazParC: Kazakh Parallel Corpus for Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Statistical machine translation gained ground over rule-based machine translation in the late 1990s thanks to its ability to learn from large bilingual corpora.
Approach: They propose to develop a parallel corpus for machine translation across Kazakh, English, Russian, and Turkish.
Outcome: The proposed model outperforms Google Translate and Yandex Translate in terms of performance and evaluation metrics.
KazQAD: Kazakh Open-Domain Question Answering Dataset (2024.lrec-main)

Copied to clipboard

Challenge: KazQAD contains just under 6,000 unique questions with extracted short answers and nearly 12,000 passage-level relevance judgements.
Approach: They introduce a Kazakh open-domain question answering dataset that can be used in reading comprehension and full ODQA settings.
Outcome: The proposed dataset can be used in reading comprehension and full ODQA settings, as well as for information retrieval experiments.
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages (2025.acl-long)

Copied to clipboard

Challenge: preparing native language MMLU benchmarks is costly and limits representativeness of evaluation datasets.
Approach: They propose to use a Turkic language MMLU benchmark to assess massive multitask language understanding capabilities.
Outcome: The proposed benchmarks are based on a Turkic language morphosyntactic and cultural benchmark . the benchmarks evaluate a diverse range of open and proprietary multilingual large language models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations