Papers with M-BERT

7 papers
A Multilingual Reading Comprehension System for more than 100 Languages (2020.coling-demos)

Copied to clipboard

Challenge: Recent advances in open domain question answering (QA) have focused on machine reading comprehension (MRC)
Approach: They propose a multilingual machine reading comprehension (MRC) demo which can answer questions in over 100 languages.
Outcome: The proposed system can answer questions in over 100 languages and integrates with IBM Watson's machine translation widget to improve language accessibility.
Scalable Cross-lingual Treebank Synthesis for Improved Production Dependency Parsers (2020.coling-industry)

Copied to clipboard

Challenge: scalable Universal Dependency (UD) treebank synthesis techniques are used to improve production-grade parsers.
Approach: They propose a data augmentation technique that uses synthetic treebanks to improve production-grade parsers.
Outcome: The proposed technique improves LAS performance on seven languages by up to two points on production models trained on original UD treebanks.
Extending Multilingual BERT to Low-Resource Languages (2020.findings-emnlp)

Copied to clipboard

Challenge: Multilingual BERT (M-BERT) has been a huge success in both supervised and zero-shot cross-lingual transfer learning.
Approach: They propose a simple but effective approach to extend multilingual BERT to any new language and show an increase in F1 on M-BERT and new languages.
Outcome: The proposed approach improves on languages already in M-BERT and out of it on other languages.
X-METRA-ADA: Cross-lingual Meta-Transfer learning Adaptation to Natural Language Understanding and Question Answering (2021.naacl-main)

Copied to clipboard

Challenge: Multilingual models have gained popularity for their zero-shot cross-lingual transfer learning capabilities, but their generalization ability is inconsistent for typologically diverse languages.
Approach: They propose a meta-learning approach that adapts MAML to learn to adapt to new languages . they extensively evaluate two cross-lingual NLU tasks using English as source and spanish as target .
Outcome: The proposed approach outperforms naive fine-tuning on cross-lingual tasks for most languages.
How Multilingual is Multilingual BERT? (P19-1)

Copied to clipboard

Challenge: Existing studies have shown that deep, contextualized language models can encode syntactic and named entity information, but they have focused on what models trained on English capture about English.
Approach: They propose a multilingual model pre-trained from monolingual Wikipedia corpora . they show that multilingual BERT is surprisingly good at zero-shot cross-lingual model transfer .
Outcome: The proposed model can find translation pairs, but it exhibits systematic deficiencies affecting certain language pairs.
BERTGen: Multi-task Generation through BERT (2021.acl-long)

Copied to clipboard

Challenge: Recent work in unsupervised and self-supervised pre-training has revolutionised the field of natural language understanding (NLU).
Approach: They propose to use multimodal and multilingual pre-trained models to extend BERT by fusing them together for language generation tasks.
Outcome: The proposed model outperforms baseline models in image captioning, machine translation and multimodal machine translation tasks and is competitive with supervised counterparts.
mLongT5: A Multilingual and Efficient Text-To-Text Transformer for Longer Sequences (2023.findings-emnlp)

Copied to clipboard

Challenge: a new text-to-text transformer is suitable for multilingual inputs . many of the current models are English-only, making them inapplicable to other languages.
Approach: They propose to extend a multilingual text-to-text transformer to handle long inputs . they use the mC4 dataset to pretrain the model to handle multilingual data .
Outcome: The proposed model performs well on multilingual summarization and question-answering tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations