Papers by Joseph Attieh
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki (2024.eacl-demo)
Copied to clipboard
Timothee Mickus, Stig-Arne Grönroos, Joseph Attieh, Michele Boggia, Ona De Gibert, Shaoxiong Ji, Niki Andreas Loppi, Alessandro Raganato, Raúl Vázquez, Jörg Tiedemann
| Challenge: | a growing trend towards modularization is limiting the size and information that can be handled in large language models. |
| Approach: | They propose a framework for training massively multilingual modular machine translation systems at scale. |
| Outcome: | The proposed framework is adapted to train multilingual models at scale on NVIDIA GPUs. |
Isotropy, Clusters, and Classifiers (2024.acl-short)
Copied to clipboard
| Challenge: | Existing evidence supports and challenges the use of isotropy in embedding spaces. |
| Approach: | They propose to formalize this connection mathematically and empirically and prove it's true . they argue that isotropy imposes requirements on embedding space that are not compatible with clusters . |
| Outcome: | The proposed method sheds light on previous studies focusing on anisotropy in embedding spaces. |
Scaling Low-Resource MT via Synthetic Data Generation with LLMs (2025.emnlp-main)
Copied to clipboard
Ona de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Zihao Li, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann
| Challenge: | a recent study has shown that LLM-generated synthetic data can improve low-resource machine translation performance . traditional data augmentation techniques like back-translation preserve the human-written target and synthesize the other . |
| Approach: | They construct a document-level synthetic corpus from English Europarl and extend it via pivoting to 147 additional language pairs. |
| Outcome: | The proposed model can significantly improve low-resource machine translation performance even when noisy. |
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models (2025.emnlp-demos)
Copied to clipboard
Hengyu Luo, Zihao Li, Joseph Attieh, Sawal Devkota, Ona de Gibert, Xu Huang, Shaoxiong Ji, Peiqin Lin, Bhavani Sai Praneeth Varma Mantina, Ananda Sreenidhi, Raúl Vázquez, Mengjie Wang, Samea Yusofi, Fei Yuan, Jörg Tiedemann
| Challenge: | Existing evaluation frameworks focus on English and a handful of high-resource languages, thereby overlooking the realistic performance of large language models in multilingual and lower-resourced scenarios. |
| Approach: | They propose a unified and lightweight framework that integrates 27 benchmarks under a standard ISO 639-3 language identifier system to enable seamless incorporation of new benchmarks. |
| Outcome: | The proposed framework integrates 27 benchmarks under a standard ISO 639-3 language identifier system, allowing for seamless incorporation of new benchmarks. |