Challenge: Masked language modeling (MLM) is a widely used self-supervised pretraining objective.
Approach: They propose to use a mask-based objective to predict a token that is replaced with a masked token given its context.
Outcome: The proposed objectives show that they should have half the complexity needed to perform comparably to MLM.

Similar Papers

Frustratingly Simple Pretraining Alternatives to Masked Language Modeling (2021.emnlp-main)

Copied to clipboard

Challenge: Masked language modeling (MLM) is widely used in natural language processing for self-supervised learning of text representations.
Approach: They propose to use token-level classification tasks as main pretraining objectives instead of Masked language modeling (MLM) . Empirical results show that pretraining a model with 41% of the BERT-BASE’s parameters, BERT MEDIUM results in only a 1% drop in GLUE scores with their best objective.
Outcome: Empirical results show that the proposed methods achieve comparable or better performance to MLM using a BERT-BASE architecture.
To Pretrain or Not to Pretrain: Examining the Benefits of Pretrainng on Resource Rich Tasks (2020.acl-main)

Copied to clipboard

Challenge: Existing studies on pretraining NLP models with variants of Masked Language Model (MLM) objectives have shown that the number of training samples used in the downstream task is limited.
Approach: They propose to use MLM objectives to pretrain NLP models with variants of Masked Language Model (MLM) objectives to improve accuracy on downstream tasks.
Outcome: The proposed model can reach a diminishing return point as the supervised data size increases significantly.
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token (2022.findings-emnlp)

Copied to clipboard

Challenge: Large-scale pre-trained MLMs can be used to generalize well to a wide range of tasks.
Approach: They propose to append [MASK]s at a later layer to reduce sequence length for earlier layers.
Outcome: The proposed method outperforms RoBERTa for 6 out of 8 GLUE tasks on average by 0.4%.
How does the pre-training objective affect what large language models learn about linguistic properties? (2022.acl-short)

Copied to clipboard

Challenge: Several pre-training objectives have been proposed to pre-train language models . but, to our knowledge, no studies have investigated how different pre- training objectives affect what BERT learns about linguistic properties.
Approach: They propose to use masked language modeling to pre-train language models . they propose to optimize a mangled language modeling objective to learn linguistic information .
Outcome: The proposed objectives improve BERT's learning of linguistic properties compared to non-linguistically motivated objectives.
A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Various types of social biases have been reported with pretrained Masked Language Models (MLMs) in prior work.
Approach: They conduct a comprehensive study on 39 pretrained MLMs to examine their model factors and their social biases.
Outcome: The proposed model factors influence social biases learned by an MLM and their downstream task performance.
Should You Mask 15% in Masked Language Modeling? (2023.eacl-main)

Copied to clipboard

Challenge: Masked language models (MLMs) traditionally mask 15% of tokens due to the belief that more masking would leave insufficient context to learn good representations.
Approach: They revisit the 15% masking rate of MLMs to examine the role of masking in linguistic training.
Outcome: The proposed masking rate outperforms BERT-large size models on GLUE and SQUAD while maintaining 95% accuracy.
Efficient Pre-training of Masked Language Model via Concept-based Curriculum Masking (2022.emnlp-main)

Copied to clipboard

Challenge: Masked language modeling (MLM) has been widely used for pre-training effective bidirectional representations but comes at a substantial training cost.
Approach: They propose a concept-based curriculum masking method that evaluates the MLM difficulty of each token based on a carefully-designed linguistic difficulty criterion.
Outcome: The proposed method significantly improves pre-training efficiency with the original BERT model at half the training cost.
On the Influence of Masking Policies in Intermediate Pre-training (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that inserting an intermediate pre-training stage improves performance of masked language models.
Approach: They propose methods to automate the discovery of optimal masking policies via direct supervision or meta-learning.
Outcome: The proposed method outperforms the heuristic of masking named entities on TriviaQA and can be generalizable beyond that task.
Data Efficient Masked Language Modeling for Vision and Language (2021.findings-emnlp)

Copied to clipboard

Challenge: Masked language modeling (MLM) is one of the key sub-tasks in vision-language pretraining.
Approach: They propose a masking strategy that masks tokens with a 15% probability for text-only data.
Outcome: The proposed masking strategy outperforms the baseline model on a prompt-based probing task designed to elicit image objects.
Difference-Masking: Choosing What to Mask in Continued Pretraining (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to masked prediction have shown that deciding what to mask can substantially improve learning outcomes.
Approach: They propose a masking strategy that automatically chooses what to mask during continued pretraining by considering what makes a task domain different from the pretraining domain.
Outcome: The proposed masking strategy outperforms baselines on language-only and multimodal video tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations