Papers by Hugo Abonizio
MonoByte: A Pool of Monolingual Byte-level Language Models (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies have shown that multilingual models can achieve zero-shot cross-lingual performance on various NLP tasks, but due to the cost of pretraining, they often use public models with limited budgets. |
| Approach: | They propose to use tokenized models to test cross-lingual ability in multilingual and monolingual corpora. |
| Outcome: | The results show that models pretrained on multilingual and even monolingual corpora perform better than models pre-trained on SOTA models. |