Papers by Diana Abagyan
Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects (2024.emnlp-main)
Copied to clipboard
Orevaoghene Ahia, Anuoluwapo Aremu, Diana Abagyan, Hila Gonen, David Adelani, Daud Abolade, Noah Smith, Yulia Tsvetkov
| Challenge: | Recent efforts to develop NLP tools for low-resource languages focus on their standard dialects. |
| Approach: | They propose a high-quality parallel text and speech corpus for Yoruba . they use native speakers to collect data from four regional yoruba dialects . |
| Outcome: | The proposed dataset shows that dialect-adaptive finetuning can narrow performance disparities . the dataset will be released publicly under an open license . |
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers (2026.acl-long)
Copied to clipboard
Diana Abagyan, Alejandro R. Salamanca, Andres Felipe Cruz-Salinas, Kris Cao, Hangyu Lin, Acyr Locatelli, Marzieh Fadaee, Ahmet Üstün, Sara Hooker
| Challenge: | Existing approaches to train multilingual large language models for many languages at once are limited due to limited model capacity, scarce high-quality data, and compute constraints. |
| Approach: | They propose to use a universal tokenizer to improve language plasticity and adaptability to new languages by up to 20%. |
| Outcome: | The proposed tokenizer improves language plasticity and improves plasticity towards languages that are completely unseen in the tokenizer and pretraining, by up to 5% win rate gain. |