Papers by Diana Abagyan

2 papers
Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects (2024.emnlp-main)

Copied to clipboard

Challenge: Recent efforts to develop NLP tools for low-resource languages focus on their standard dialects.
Approach: They propose a high-quality parallel text and speech corpus for Yoruba . they use native speakers to collect data from four regional yoruba dialects .
Outcome: The proposed dataset shows that dialect-adaptive finetuning can narrow performance disparities . the dataset will be released publicly under an open license .
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to train multilingual large language models for many languages at once are limited due to limited model capacity, scarce high-quality data, and compute constraints.
Approach: They propose to use a universal tokenizer to improve language plasticity and adaptability to new languages by up to 20%.
Outcome: The proposed tokenizer improves language plasticity and improves plasticity towards languages that are completely unseen in the tokenizer and pretraining, by up to 5% win rate gain.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations