Papers by Joseph Attieh

4 papers
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki (2024.eacl-demo)

Copied to clipboard

Challenge: a growing trend towards modularization is limiting the size and information that can be handled in large language models.
Approach: They propose a framework for training massively multilingual modular machine translation systems at scale.
Outcome: The proposed framework is adapted to train multilingual models at scale on NVIDIA GPUs.
Isotropy, Clusters, and Classifiers (2024.acl-short)

Copied to clipboard

Challenge: Existing evidence supports and challenges the use of isotropy in embedding spaces.
Approach: They propose to formalize this connection mathematically and empirically and prove it's true . they argue that isotropy imposes requirements on embedding space that are not compatible with clusters .
Outcome: The proposed method sheds light on previous studies focusing on anisotropy in embedding spaces.
Scaling Low-Resource MT via Synthetic Data Generation with LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that LLM-generated synthetic data can improve low-resource machine translation performance . traditional data augmentation techniques like back-translation preserve the human-written target and synthesize the other .
Approach: They construct a document-level synthetic corpus from English Europarl and extend it via pivoting to 147 additional language pairs.
Outcome: The proposed model can significantly improve low-resource machine translation performance even when noisy.
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on English and a handful of high-resource languages, thereby overlooking the realistic performance of large language models in multilingual and lower-resourced scenarios.
Approach: They propose a unified and lightweight framework that integrates 27 benchmarks under a standard ISO 639-3 language identifier system to enable seamless incorporation of new benchmarks.
Outcome: The proposed framework integrates 27 benchmarks under a standard ISO 639-3 language identifier system, allowing for seamless incorporation of new benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations