Papers by Raúl Vázquez

8 papers
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki (2024.eacl-demo)

Copied to clipboard

Challenge: a growing trend towards modularization is limiting the size and information that can be handled in large language models.
Approach: They propose a framework for training massively multilingual modular machine translation systems at scale.
Outcome: The proposed framework is adapted to train multilingual models at scale on NVIDIA GPUs.
A Closer Look at Parameter Contributions When Training Neural Language and Translation Models (2022.coling-1)

Copied to clipboard

Challenge: Neural models and Transformers have been used for almost every NLP task . however, the intrinsic dynamics of the training procedure have not been studied in depth for highly complex network architectures.
Approach: They analyze the learning dynamics of neural language and translation models using Loss Change Allocation indicator . they use a standard Transformer architecture to train a model with three learning objectives .
Outcome: The proposed model is based on a standard model that is used for training tasks.
Language Models Learn Universal Representations of Numbers and Here’s Why You Should Care (2026.acl-long)

Copied to clipboard

Challenge: Prior work has shown that large language models (LLMs) often converge to accurate input embedding for numbers, based on sinusoidal representations.
Approach: They show that large language models often converge to accurate input embedding for numbers, based on sinusoidal representations.
Outcome: The proposed representations are strikingly systematic, and are interchangeable in a large swathe of experimental setups.
An Empirical Investigation of Word Alignment Supervision for Zero-Shot Multilingual Neural Machine Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work has highlighted several flaws of MNMT models in zero-shot scenarios where language labels are ignored and the wrong language is generated.
Approach: They propose to combine explicit alignment to language labels with word alignment supervision to improve zero-shot translations.
Outcome: The proposed model improves on three multilingual MT benchmarks.
On the differences between BERT and MT encoder spaces and how to address them in translation tasks (2021.acl-srw)

Copied to clipboard

Challenge: Various studies show that pretrained language models cannot replace encoders in neural machine translation despite their success in other tasks.
Approach: They propose a supervised transformation from one into the other to improve the applicability of BERT in neural machine translation.
Outcome: The proposed transformations show that they cannot replace encoders in MT despite their success in other tasks.
Scaling Low-Resource MT via Synthetic Data Generation with LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that LLM-generated synthetic data can improve low-resource machine translation performance . traditional data augmentation techniques like back-translation preserve the human-written target and synthesize the other .
Approach: They construct a document-level synthetic corpus from English Europarl and extend it via pivoting to 147 additional language pairs.
Outcome: The proposed model can significantly improve low-resource machine translation performance even when noisy.
Your Model is Overconfident, and Other Lies We Tell Ourselves (2025.acl-long)

Copied to clipboard

Challenge: Analyzing 29 models, we find that difficulty is not linear or monotonic.
Approach: They examine the interplay and divergence among various metrics for assessing intrinsic difficulty, including annotator dissensus, training dynamics, and model confidence.
Outcome: The proposed model is based on 29 models on three datasets and analyzed by a linguistics team.
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on English and a handful of high-resource languages, thereby overlooking the realistic performance of large language models in multilingual and lower-resourced scenarios.
Approach: They propose a unified and lightweight framework that integrates 27 benchmarks under a standard ISO 639-3 language identifier system to enable seamless incorporation of new benchmarks.
Outcome: The proposed framework integrates 27 benchmarks under a standard ISO 639-3 language identifier system, allowing for seamless incorporation of new benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations