Papers by Marc’Aurelio Ranzato

8 papers
Facebook AI’s WAT19 Myanmar-English Translation Task Submission (D19-52)

Copied to clipboard

Challenge: Using back-translation, we can improve generalization by using noisy channel re-ranking and ensembling.
Approach: They propose to use BPE-based transformer models to leverage monolingual data to improve generalization and use noisy channel re-ranking and ensembling to improve results.
Outcome: The proposed system improves on the baseline system trained exclusively on the provided small parallel dataset, and the human evaluation and BLEU score are higher.
The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali–English and Sinhala–English (D19-1)

Copied to clipboard

Challenge: a vast majority of language pairs in the world are considered low-resource because they have little parallel data available.
Approach: They propose to use a dataset to evaluate methods trained on low-resource language pairs . they report baseline performance using supervised, weakly supervised and semi-supervised settings .
Outcome: The proposed evaluation datasets show that current state-of-the-art methods perform poorly on this benchmark, posing a challenge to the research community working on low-resource MT.
The Flores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation (2022.tacl-1)

Copied to clipboard

Challenge: a lack of good evaluation benchmarks hinders progress in low-resource and multilingual machine translation . despite advances in translation quality for a handful of languages, many low-source languages are not even supported by most popular translation engines.
Approach: They propose a high-quality evaluation benchmark for machine translation using 3001 sentences from Wikipedia . they aim to improve evaluation of models on long tail of low-resource languages .
Outcome: The proposed evaluation benchmarks are based on 3001 sentences extracted from Wikipedia . the results show that the models can be used to evaluate multilingual systems .
Phrase-Based & Neural Unsupervised Machine Translation (D18-1)

Copied to clipboard

Challenge: Recent advances in machine translation have reported near human-level performance on several languages, yet their effectiveness strongly relies on the availability of large amounts of parallel sentences.
Approach: They propose two models that leverage a careful initialization of the parameters and denoising effect of language models.
Outcome: The proposed models outperform the current methods on English-French and German-English benchmarks while being simpler and having fewer hyper-parameters.
On The Evaluation of Machine Translation Systems Trained With Back-Translation (2020.acl-main)

Copied to clipboard

Challenge: Back-translation is a data augmentation technique that can be used to improve neural machine translation systems.
Approach: They propose to combine back-translation with a language model score to measure fluency.
Outcome: The proposed method improves translation quality of natural text and translationese according to professional translators.
Classical Structured Prediction Losses for Sequence to Sequence Learning (N18-1)

Copied to clipboard

Challenge: Recent work on training neural attention models at the sequence level has focused on a series of objective functions commonly used for structured prediction.
Approach: They propose to use objective functions commonly used to train linear models for structured prediction to train neural attention models at the sequence-level using either reinforcement learning-style methods or beam search optimization.
Outcome: The proposed model outperforms beam search optimization on German-English translation and abstractive summarization tasks.
Discriminative Reranking for Neural Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: reranking models allow the integration of rich features to select a better output hypothesis within an n-best list or lattice.
Approach: They use discriminative reranking to train a large transformer architecture to train an ranked list of hypotheses.
Outcome: Experiments on four WMT directions show that discriminative reranking improves translation quality.
The Source-Target Domain Mismatch Problem in Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Despite the interconnected world we live in, people in different places talk about different things in different parts of the world.
Approach: They propose a metric to quantify the effect of local context in machine translation and propose measurable results.
Outcome: The proposed metric can be used to quantify the effect of local context on the use of language in machine translation systems on low resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations