Papers with Turkic
Visual Modeling of Turkish Morphology (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, there are three publicly accessible morphological analyzers for Turkish . |
| Approach: | They propose to make modeling easier and more maintainable by using diagramming tools and automating much of the code generation. |
| Outcome: | The proposed model can be easily maintained and the code generation automated. |
Cross-Lingual Word Embeddings for Turkic Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing techniques to align monolingual embeddings are difficult to use because of low resources. |
| Approach: | They propose to use existing techniques to align monolingual embedding spaces for Turkic, Uzbek, Azeri, Kazakh and Kyrgyz languages. |
| Outcome: | The proposed techniques outperform existing techniques on bilingual dictionaries and an extrinsic task. |
KazakhTTS2: Extending the Open-Source Kazakh TTS Corpus With More Data, Speakers, and Topics (2022.lrec-1)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) is a process of converting written text into speech. |
| Approach: | They present an expanded version of their text-to-speech corpus for Kazakh . they propose to use the corpus to build high-quality TTS systems for the language . |
| Outcome: | The constructed corpus is sufficient to build robust TTS models for Kazakh and other Turkic languages, with a subjective mean opinion score ranging from 3.6 to 4.2 for all the five speakers. |
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages (2025.acl-long)
Copied to clipboard
Jafar Isbarov, Arofat Akhundjanova, Mammad Hajili, Kavsar Huseynova, Dmitry Gaynullin, Anar Rzayev, Osman Tursun, Aizirek Turdubaeva, Ilshat Saetov, Rinat Kharisov, Saule Belginova, Ariana Kenbayeva, Amina Alisheva, Abdullatif Köksal, Samir Rustamov, Duygu Ataman
| Challenge: | preparing native language MMLU benchmarks is costly and limits representativeness of evaluation datasets. |
| Approach: | They propose to use a Turkic language MMLU benchmark to assess massive multitask language understanding capabilities. |
| Outcome: | The proposed benchmarks are based on a Turkic language morphosyntactic and cultural benchmark . the benchmarks evaluate a diverse range of open and proprietary multilingual large language models . |