Papers by Rami Al-Rfou

7 papers
nmT5 - Is parallel data still relevant for pre-training massively multilingual language models? (2021.acl-short)

Copied to clipboard

Challenge: Recent studies have shown that cross-lingual transfer learning in pre-trained multilingual models could be improved further by incorporating parallel data.
Approach: They propose to integrate parallel data into mT5 pre-training to improve results on downstream multilingual and cross-lingual tasks.
Outcome: The proposed model improves cross-lingual transfer significantly in small fine-tuning datasets and small model sizes.
LAReQA: Language-Agnostic Answer Retrieval from a Multilingual Pool (2020.emnlp-main)

Copied to clipboard

Challenge: LAReQA tests for “strong” cross-lingual alignment, requiring semantically related cross-language pairs to be closer in representation space than unrelated same-language pair.
Approach: They propose a new benchmark for language-agnostic answer retrieval from a multilingual candidate pool that tests for "strong" cross-lingual alignment . they augment training data via machine translation and find that model performance is improved by augmenting training data through machine translation .
Outcome: The proposed task is based on multilingual BERT (mBERT) and XLM-R.
The Power of Scale for Parameter-Efficient Prompt Tuning (2021.emnlp-main)

Copied to clipboard

Challenge: Unlike discrete text prompts used by GPT-3, soft prompts are learned through backpropagation and can be tuned to incorporate signals from any number of labeled examples.
Approach: They propose a mechanism for learning "soft prompts" to condition frozen language models to perform specific downstream tasks.
Outcome: The proposed method outperforms fewshot learning using GPT-3 and matches the quality of model tuning as models exceed billions of parameters.
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models (2022.tacl-1)

Copied to clipboard

Challenge: a number of pre-trained language models use sequences of tokens corresponding to word units . token-free models that operate directly on raw text have many advantages .
Approach: They propose a standard Transformer architecture that can be used to process byte sequences . they also characterize trade-offs in terms of parameter count, training FLOPs, and inference speed .
Outcome: The proposed model is more robust to noise and more robust on spelling and pronunciation tasks.
Wiki-40B: Multilingual Language Model Dataset (2020.lrec-1)

Copied to clipboard

Challenge: We propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families.
Approach: They propose a multilingual language model benchmark composed of 40+ languages . they train monolingual causal language models using a state-of-the-art model .
Outcome: The proposed model is composed of 40+ languages spanning several scripts and linguistic families.
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer (2021.naacl-main)

Copied to clipboard

Challenge: Current natural language processing pipelines often use transfer learning, where a model is pre-trained on a data-rich task before being fine-tuned on . this significantly limits their use given that roughly 80% of the world population does not speak English.
Approach: They introduce a multilingual variant of T5 that was pre-trained on a new Common Crawl-based dataset covering 101 languages.
Outcome: The proposed model achieves state-of-the-art on multilingual benchmarks and a simple technique to prevent accidental translation in the zero-shot setting.
Knowledge Graph Based Synthetic Corpus Generation for Knowledge-Enhanced Language Model Pre-training (2021.naacl-main)

Copied to clipboard

Challenge: Existing work on data-to-text generation focused on domain-specific benchmark datasets.
Approach: They use a KG-Wikipedia text aligned corpus to verbalize the entire English Wikidata KG . they show that this approach can be used to integrate structured KGs and natural language corpora .
Outcome: The proposed method improves on open domain QA and the LAMA knowledge probe.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations