Papers by Clément Lefebvre

1 papers
LuxemBERT: Simple and Practical Data Augmentation in Language Model Pre-Training for Luxembourgish (2022.lrec-1)

Copied to clipboard

Challenge: Pre-trained Language Models such as BERT are ubiquitous in NLP but are scarce for low-resource languages such as Luxembourgish.
Approach: They propose a BERT model for Luxembourgish language that they use to augment pre-training datasets by partially translating text data from a closely related language.
Outcome: The proposed model outperforms the baseline model and the mBERT model in Luxembourgish.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations