Papers by Tom Vanallemeersch

3 papers
Being Generous with Sub-Words towards Small NMT Children (2020.lrec-1)

Copied to clipboard

Challenge: In the context of under-resourced neural machine translation, transfer learning from an NMT model trained on a high resource language pair, or from a multilingual NMT (M-NMT) model, has been shown to boost performance to a large extent.
Approach: They propose to use a multilingual NMT model to train on an under-resourced child and to use large sub-word vocabularies to improve performance.
Outcome: The proposed approach involving dynamic vocabularies is both practical and effective on two under-resourced language pairs, i.e. Icelandic-English and Irish-English.
A Post-Editing Dataset in the Legal Domain: Do we Underestimate Neural Machine Translation Quality? (2020.lrec-1)

Copied to clipboard

Challenge: Current state-of-the-art in Neural Machine Translation (NMT) has reached remarkable progress, but human evaluations are often judged as having lower quality than top NMT systems.
Approach: They propose to use a machine translation dataset with post-edited high-quality neural machine translation and independent human references to compare the results.
Outcome: The proposed dataset includes 31K tuples including a source sentence, the respective machine translation by a neural machine translation system, and a post-edited version of such translation by professional translator.
ELRC Action: Covering Confidentiality, Correctness and Cross-linguality (2022.lrec-1)

Copied to clipboard

Challenge: ELRC aims to reduce language barriers by assessing language technology (LT) specifications . automated anonymisation and multilingual fake news processing are two of the most extensive LT assessments .
Approach: They describe language technology (LT) assessments carried out by the European Commission . they zoom in on two of the most extensive assessments, namely automated anonymisation and multilingual fake news processing.
Outcome: The language technology (LT) assessments carried out by the European Commission are detailed in this paper . they include a consultation round with stakeholders from public organisations, academia and industry . the ELRC action aims to create proof-of-concept environments integrating relevant tools and services .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations