Papers by Francisco Guzman
Small Data, Big Impact: Leveraging Minimal Data for Effective Machine Translation (2023.acl-long)
Copied to clipboard
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzman
| Challenge: | Existing datasets are not economical to create large-scale datasets, but for low-resource languages, a few thousand professionally translated sentence pairs can be useful. |
| Approach: | They propose to use a dataset to train machine translation models on pre-existing and synthetic data to augment them with millions of sentences through backtranslation. |
| Outcome: | The proposed model can cover hundreds of languages with high quality training data even when smaller but lower quality datasets are used. |