Papers by Hasti Toossi

2 papers
Predicting Machine Translation Performance on Low-Resource Languages: The Role of Domain Similarity (2024.findings-eacl)

Copied to clipboard

Challenge: Existing approaches for predicting the performance of NLP models for low-resource languages (LRLs) focus on high-resourced languages, overlooking LRLs and domain shifts.
Approach: They investigate the impact of domain similarity on predicting performance of machine translation models in low-resource languages.
Outcome: The results show that domain similarity has the most important impact on predicting the performance of Machine Translation models.
A Reproducibility Study on Quantifying Language Similarity: The Impact of Missing Values in the URIEL Knowledge Base (2024.naacl-srw)

Copied to clipboard

Challenge: URIEL aggregates linguistic information for 4,005 languages and computes distances based on this information.
Approach: They propose to use a typological knowledge base to quantify language similarity to investigate URIEL's ambiguity in calculating language distances and handling missing values.
Outcome: The URIEL knowledge base does not provide information about typological features for 31% of the languages it represents, undermining the reliability of the database, especially on low-resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations