Papers by Tom Vanallemeersch
Being Generous with Sub-Words towards Small NMT Children (2020.lrec-1)
Copied to clipboard
| Challenge: | In the context of under-resourced neural machine translation, transfer learning from an NMT model trained on a high resource language pair, or from a multilingual NMT (M-NMT) model, has been shown to boost performance to a large extent. |
| Approach: | They propose to use a multilingual NMT model to train on an under-resourced child and to use large sub-word vocabularies to improve performance. |
| Outcome: | The proposed approach involving dynamic vocabularies is both practical and effective on two under-resourced language pairs, i.e. Icelandic-English and Irish-English. |
A Post-Editing Dataset in the Legal Domain: Do we Underestimate Neural Machine Translation Quality? (2020.lrec-1)
Copied to clipboard
Julia Ive, Lucia Specia, Sara Szoc, Tom Vanallemeersch, Joachim Van den Bogaert, Eduardo Farah, Christine Maroti, Artur Ventura, Maxim Khalilov
| Challenge: | Current state-of-the-art in Neural Machine Translation (NMT) has reached remarkable progress, but human evaluations are often judged as having lower quality than top NMT systems. |
| Approach: | They propose to use a machine translation dataset with post-edited high-quality neural machine translation and independent human references to compare the results. |
| Outcome: | The proposed dataset includes 31K tuples including a source sentence, the respective machine translation by a neural machine translation system, and a post-edited version of such translation by professional translator. |
ELRC Action: Covering Confidentiality, Correctness and Cross-linguality (2022.lrec-1)
Copied to clipboard
Tom Vanallemeersch, Arne Defauw, Sara Szoc, Alina Kramchaninova, Joachim Van den Bogaert, Andrea Lösch
| Challenge: | ELRC aims to reduce language barriers by assessing language technology (LT) specifications . automated anonymisation and multilingual fake news processing are two of the most extensive LT assessments . |
| Approach: | They describe language technology (LT) assessments carried out by the European Commission . they zoom in on two of the most extensive assessments, namely automated anonymisation and multilingual fake news processing. |
| Outcome: | The language technology (LT) assessments carried out by the European Commission are detailed in this paper . they include a consultation round with stakeholders from public organisations, academia and industry . the ELRC action aims to create proof-of-concept environments integrating relevant tools and services . |