Papers by Teemu Vahtola
Modeling Noise in Paraphrase Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | Noisy labels in training data are challenging and can lead to incorrect decisions . large pre-trained language models have achieved great results in many NLP tasks . |
| Approach: | They propose to use a linear noise model to augment pre-trained language models to account for label noise in fine-tuning. |
| Outcome: | The proposed model can be applied without further knowledge about annotation quality and label confidence of training examples and their results are compared with other models. |
Guiding Zero-Shot Paraphrase Generation with Fine-Grained Control Tokens (2023.starsem-1)
Copied to clipboard
| Challenge: | Sequence-to-sequence paraphrase generation models struggle with the generation of diverse paraphrases. |
| Approach: | They propose a translation-based guided paraphrase generation model that learns useful features for promoting surface form variation in generated paraphrases from cross-lingual parallel data. |
| Outcome: | The proposed model learns useful features for promoting surface form variation in generated paraphrases from cross-lingual parallel data. |
Scaling Low-Resource MT via Synthetic Data Generation with LLMs (2025.emnlp-main)
Copied to clipboard
Ona de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Zihao Li, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann
| Challenge: | a recent study has shown that LLM-generated synthetic data can improve low-resource machine translation performance . traditional data augmentation techniques like back-translation preserve the human-written target and synthesize the other . |
| Approach: | They construct a document-level synthetic corpus from English Europarl and extend it via pivoting to 147 additional language pairs. |
| Outcome: | The proposed model can significantly improve low-resource machine translation performance even when noisy. |