UltraSparseBERT: 99% Conditionally Sparse Language Modelling (2024.acl-short)

Copied to clipboard

Challenge: UltraSparseBERT uses 0.3% of its neurons during inference, compared to similar BERT models.
Approach: They propose a BERT variant that selectively engages just 12 out of 4095 neurons for each layer inference.
Outcome: The proposed model employs 0.3% of its neurons during inference while performing on par with similar models.

Similar Papers

SkipBERT: Efficient Inference with Shallow Layer Skipping (2022.acl-long)

Copied to clipboard

Challenge: Pre-trained language models have significant demands in computation and inference time, limiting their use in resource-constrained or latencysensitive applications.
Approach: They propose to encode text chunks into independent representations and skip computation of shallow layers to accelerate inference.
Outcome: The proposed approach can reduce latency by 65% without sacrificing performance.
EfficientBERT: Progressively Searching Multilayer Perceptron via Warm-up Knowledge Distillation (2021.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models have shown remarkable results on various NLP tasks.
Approach: They propose to improve the feed-forward network (FFN) in BERT with a higher computational cost than improving the multi-head attention (MHA).
Outcome: The proposed model is 6.9 smaller and 4.4 faster than BERTBASE and has competitive performances on GLUE and SQuAD Benchmarks.
KinyaBERT: a Morphology-aware Kinyarwanda Language Model (2022.acl-long)

Copied to clipboard

Challenge: Pre-trained language models such as BERT are sub-optimal at handling morphologically rich languages.
Approach: They propose a two-tier BERT architecture that leverages a morphological analyzer and explicitly represents morphology in a low-resource Kinyarwanda language.
Outcome: The proposed model outperforms baseline models on the low-resource morphologically rich Kinyarwanda language by 2% in F1 score and 4.3% in average score of GLUE benchmark.
The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Pre-trained Transformer models provide robust language representations which can be specialized on various tasks.
Approach: They propose an efficient pruning method based on approximate second-order information that allows pruning weight blocks to be used for pruning.
Outcome: The proposed method is the first to be applied at the BERT scale and significantly pushes the boundaries of the current sparse models with respect to all metrics: model size, inference speed and task accuracy.
TernaryBERT: Distillation-aware Ultra-low Bit BERT (2020.emnlp-main)

Copied to clipboard

Challenge: Transformer-based pre-training models like BERT are computationally expensive and limited to resource-constrained devices.
Approach: They propose a method which ternarizes the weights in a fine-tuned BERT model.
Outcome: The proposed method outperforms the other methods on the GLUE and SQUAD benchmarks while being 14.9x smaller.
HybridBERT - Making BERT Pretraining More Efficient Through Hybrid Mixture of Attention Mechanisms (2024.naacl-srw)

Copied to clipboard

Challenge: Pretrained transformer-based language models have produced state-of-the-art performance in most natural language understanding tasks.
Approach: They propose two hybrid architectures that combine self-attention and additive attention mechanisms with sub-layer normalization to achieve double the pretraining accuracy of a vanilla-BERT baseline.
Outcome: The proposed architectures outperform BERT-base on two downstream tasks while accelerating inference.
Fast and Accurate Deep Bidirectional Language Representations for Unsupervised Learning (2020.acl-main)

Copied to clipboard

Challenge: Existing deep bidirectional language models are limited by repetitive inferences on unsupervised tasks for the computation of contextual language representations.
Approach: They propose a deep bidirectional language model called a Transformer-based Text Autoencoder (T-TA) it computes contextual language representations without repetition and shows competitive or even better accuracies than BERT .
Outcome: The proposed model performs six times faster on a reranking task and twelve times faster in a semantic similarity task.
Blockwise Self-Attention for Long Document Understanding (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in pre-training and fine-tuning methods have drastically reshaped the landscape of natural language processing research.
Approach: They propose a lightweight BERT model that introduces sparse block structures into the attention matrix to reduce memory consumption and training/inference time.
Outcome: The proposed model uses 18.7-36.1% less memory and 12.0-25.1% more time to learn compared to an advanced BERT-based model, RoBERTa.
DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference (2020.acl-main)

Copied to clipboard

Challenge: Large-scale pre-trained language models such as BERT are notorious for being slow in both training and inference.
Approach: They propose a method to accelerate BERT inference by inserting extra classification layers between each transformer layer of BERT.
Outcome: The proposed method saves up to 40% inference time with minimal degradation in model quality.
TR-BERT: Dynamic Token Reduction for Accelerating BERT Inference (2021.naacl-main)

Copied to clipboard

Challenge: Existing pre-trained language models (PLMs) are expensive in inference, making them impractical in resource-limited real-world applications.
Approach: They propose a dynamic token reduction approach to accelerate PLMs' inference by adapting the layer number of each token to avoid redundant calculation.
Outcome: The proposed approach speeds up BERT by 2-5 times and improves performance in long-text tasks with less computation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations