Multi-armed bandits for resource efficient, online optimization of language model pre-training: the use case of dynamic masking (2023.findings-acl)
Copied to clipboard
| Challenge: | Using a Bayesian optimization framework, we pre-train Transformer-based language models (TLMs) using a multi-armed bandit framework requires high computational resources and introduces many unresolved design choices. |
| Approach: | They propose a Bayesian optimization framework for resource efficient pre-training of Transformer-based language models. |
| Outcome: | The proposed framework achieves lower MLM loss in fewer epochs, across settings, while avoiding expensive hyperparameter grid search. |
Similar Papers
Large Language Model-Enhanced Multi-Armed Bandits (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have been used to sequential decision-making tasks like multi-armed bandits where an LLM is tasked with selecting arms in each iteration is often suboptimal. |
| Approach: | They propose to combine MAB and LLMs to leverage the in-context learning capability of LLM for reward prediction. |
| Outcome: | The proposed approach outperforms LLM-based direct arm selection on synthetic tasks where only preference feedback between arm pairs is available. |
CDLM: Cross-Document Language Modeling (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing language models (LMs) provide powerful representations for internal text structure, but there are important applications for multi-text tasks. |
| Approach: | They propose a pretraining approach that incorporates two key ideas into the masked language modeling objective. |
| Outcome: | The proposed model improves over existing models and sets of long-range transformers and can be easily applied to multiple multi-text tasks. |
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale pre-trained MLMs can be used to generalize well to a wide range of tasks. |
| Approach: | They propose to append [MASK]s at a later layer to reduce sequence length for earlier layers. |
| Outcome: | The proposed method outperforms RoBERTa for 6 out of 8 GLUE tasks on average by 0.4%. |
NB-MLM: Efficient Domain Adaptation of Masked Language Models for Sentiment Analysis (2021.emnlp-main)
Copied to clipboard
| Challenge: | Pre-training Masked Language Models (MLMs) on massive datasets is expensive, but it is performed for each domain or task individually and is resource-demanding. |
| Approach: | They propose a method for more efficient adaptation that focuses on predicting words with large weights of the Naive Bayes classifier trained for the task at hand. |
| Outcome: | The proposed method improves sentiment analysis by focusing on predicting words with large weights of the Naive Bayes classifier trained for the task at hand. |
Train No Evil: Selective Masking for Task-Guided Pre-Training (2020.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained language models can't capture domain-specific and task-specific patterns because of the task-agnostic pre-training stage. |
| Approach: | They propose a task-guided pre-training stage with selective masking between general pre-train and fine-tuning to learn domain-specific patterns. |
| Outcome: | The proposed method can achieve comparable or even better performance with less than 50% of computation cost. |
Frustratingly Simple Pretraining Alternatives to Masked Language Modeling (2021.emnlp-main)
Copied to clipboard
| Challenge: | Masked language modeling (MLM) is widely used in natural language processing for self-supervised learning of text representations. |
| Approach: | They propose to use token-level classification tasks as main pretraining objectives instead of Masked language modeling (MLM) . Empirical results show that pretraining a model with 41% of the BERT-BASE’s parameters, BERT MEDIUM results in only a 1% drop in GLUE scores with their best objective. |
| Outcome: | Empirical results show that the proposed methods achieve comparable or better performance to MLM using a BERT-BASE architecture. |
Inexpensive Domain Adaptation of Pretrained Language Models: Case Studies on Biomedical NER and Covid-19 QA (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Pretrained Language Models (PTLMs) are typically pretraining on target-domain text, which is expensive in terms of hardware, runtime and CO 2 emissions. |
| Approach: | They propose a faster, CPU-only domainadaptation method that trains Word2Vec on target-domain text and aligns the resulting word vectors with the wordpiece vectors of a general-domain PTLM. |
| Outcome: | The proposed method covers 60% of the BioBERT - BERT F1 delta, 5% of BioBERTS’s CO2 footprint and 2% of its cloud compute cost. |
Masking as an Efficient Alternative to Finetuning for Pretrained Language Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Extensive evaluations of masking BERT, RoBERTa, and DistilBERT on eleven diverse NLP tasks show that our binary masked language models encode information necessary for solving downstream tasks. |
| Approach: | They propose an efficient method of utilizing pretrained language models where selective binary masks are learned instead of finetuning. |
| Outcome: | Extensive evaluations of masking BERT, RoBERTa, and DistilBERT on eleven diverse NLP tasks show that the proposed method yields comparable performance to finetuning, but has a much smaller memory footprint when multiple tasks need to be solved. |
Dynamic and Efficient Inference for Text Generation via BERT Family (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods to improve performance of pre-trained language models are limited due to large-scale parameters and the universal autoregressive decoding paradigm. |
| Approach: | They propose a novel fine-tuning method which can make a single pre-trained model support Dynamic and Efficient infERence and achieve an adaptive trade-off between model performance and latency. |
| Outcome: | The proposed method achieves higher BLEU scores than the strong autoregressive Transformer model on translation tasks with 3 12 times speedup and faster inference speed compared with the BART model on four GLGE benchmark tasks. |
INGENIOUS: Using Informative Data Subsets for Efficient Pre-Training of Language Models (2023.findings-emnlp)
Copied to clipboard
H S V N S Kowndinya Renduchintala, Krishnateja Killamsetty, Sumit Bhatia, Milan Aggarwal, Ganesh Ramakrishnan, Rishabh Iyer, Balaji Krishnamurthy
| Challenge: | Pre-trained language models have a remarkable improvement in generalization capability . however, this leads to prohibitively long training times and a detrimental environmental impact . |
| Approach: | They propose to use submodular optimization to select highly informative subsets of training data to train multiple PTLMs using only fractions of data. |
| Outcome: | The proposed framework achieves 99% of the performance of fully-trained models using only fraction of training data. |