Papers by Diana Turmakhan

5 papers
Qorǵau: Evaluating Safety in Kazakh-Russian Bilingual Contexts (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have the potential to generate harmful content, posing risks to users.
Approach: They propose a dataset specifically designed for safety evaluation in Kazakh and Russian . they use a bilingual context in Kazakhstan where both Kazakh (a low-resource language) and Russian (a high-resourced language)
Outcome: The proposed dataset is designed for safety evaluation in Kazakh and Russian . it shows that both multilingual and language-specific LLMs perform better than others .
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan (2025.acl-long)

Copied to clipboard

Challenge: Kazakh language remains underrepresented in the field of natural language processing despite the country's population exceeding twenty million . however, there is a lack of dedicated models and benchmark evaluations specifically tailored to Kazakh languages.
Approach: They propose to create a dataset specifically designed for Kazakh language with 23,000 questions sourced from authentic educational materials and manually validated by native speakers and educators.
Outcome: The first MMLU-style dataset specifically designed for Kazakh language.
FRAPPE: FRAming, Persuasion, and Propaganda Explorer (2024.eacl-demo)

Copied to clipboard

Challenge: FRAPPE is a linguistic analysis, persuasion, and propaganda-based news analysis system that analyzes articles for genre, framings, and persulasion techniques.
Approach: They propose a FRAming, Persuasion, and Propaganda Explorer system that analyzes articles for genre, framings, and use of persuation techniques.
Outcome: FRAPPE analyzes articles for genre, framings, and use of persuasion techniques . it also draws comparisons between persulasion and framping strategies adopted by a diverse pool of news outlets and countries across multiple languages for different topics .
Stereotype Bias in a Bilingual Setting: A Culturally Grounded Evaluation in Kazakhstan (2026.acl-long)

Copied to clipboard

Challenge: Stereotype bias in language models is largely understudied in English . language models perform strongly on downstream NLP tasks, but they are pre-trained on large text corpora .
Approach: They use a dataset to assess stereotype bias in language models in Kazakhstan . they find that stereotype bias is most pronounced in code-mixed inputs .
Outcome: The proposed dataset shows that stereotype bias is most pronounced in code-mixed inputs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations