Papers by Hasti Toossi
Predicting Machine Translation Performance on Low-Resource Languages: The Role of Domain Similarity (2024.findings-eacl)
Copied to clipboard
Eric Khiu, Hasti Toossi, Jinyu Liu, Jiaxu Li, David Anugraha, Juan Flores, Leandro Roman, A. Seza Doğruöz, En-Shiun Lee
| Challenge: | Existing approaches for predicting the performance of NLP models for low-resource languages (LRLs) focus on high-resourced languages, overlooking LRLs and domain shifts. |
| Approach: | They investigate the impact of domain similarity on predicting performance of machine translation models in low-resource languages. |
| Outcome: | The results show that domain similarity has the most important impact on predicting the performance of Machine Translation models. |
A Reproducibility Study on Quantifying Language Similarity: The Impact of Missing Values in the URIEL Knowledge Base (2024.naacl-srw)
Copied to clipboard
| Challenge: | URIEL aggregates linguistic information for 4,005 languages and computes distances based on this information. |
| Approach: | They propose to use a typological knowledge base to quantify language similarity to investigate URIEL's ambiguity in calculating language distances and handling missing values. |
| Outcome: | The URIEL knowledge base does not provide information about typological features for 31% of the languages it represents, undermining the reliability of the database, especially on low-resource languages. |