Challenge: State-of-the-art neural machine translation methods use huge amounts of parameters.
Approach: They propose an all-inclusive quantization strategy for the Transformer to reduce computational costs and improve translation quality.
Outcome: The proposed method achieves state-of-the-art results on most tasks compared to previous methods .

Similar Papers

Extremely Low Bit Transformer Quantization for On-Device Neural Machine Translation (2020.findings-emnlp)

Copied to clipboard

Challenge: Quantization is an effective technique to address heavy computation load and memory overhead during inference.
Approach: They propose a low-bit quantization strategy to represent Transformer weights by an extremely low number of bits.
Outcome: The proposed model achieves 11.8 smaller model size than baseline model, with less than -0.5 BLEU.
Understanding and Overcoming the Challenges of Efficient Transformer Quantization (2021.emnlp-main)

Copied to clipboard

Challenge: Recent advances in transformer quantization have shown remarkable improvement in many Natural Language Processing tasks and beyond.
Approach: They propose a novel quantization scheme for transformers that can be quantized to ultra-low bit-widths, leading to significant memory savings with a minimum accuracy loss.
Outcome: The proposed methods achieve state-of-the-art results on the GLUE benchmark using BERT, while preserving memory and accuracy.
On the Way to Lossless Compression of Language Transformers: Exploring Cross-Domain Properties of Quantization (2024.lrec-main)

Copied to clipboard

Challenge: Modern Natural Language Processing models have a huge capacity, but this makes it difficult to employ.
Approach: They propose a method to quantize at least 95% of Transformer weights without access to task-specific data so the drop in performance does not exceed 0.02%.
Outcome: The proposed method quantizes 95% of Transformer weights and corresponding activations to INT8 without access to task-specific data so the drop in performance does not exceed 0.02%.
Exploring Quantization for Efficient Pre-Training of Transformer Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Quantization has proven to be effective after pre-training and during fine-tuning, but its effects on pre-trainer performance have remained unexplored.
Approach: They propose a linear quantization strategy to be applied during the pre-training of Transformers to improve model efficiency and stability.
Outcome: The proposed method improves model efficiency, stability, and performance while maintaining language modeling ability.
Self-Distilled Quantization: Achieving High Compression Rates in Transformer-Based Language Models (2023.acl-short)

Copied to clipboard

Challenge: Existing methods for quantization-aware training and quantization for learning have limitations in dealing with accumulative quantization errors.
Approach: They propose a method that minimizes accumulative quantization errors and outperforms baselines by distilling knowledge from a fine-tuned teacher network.
Outcome: The proposed method minimizes accumulative quantization errors and outperforms baselines on the XGLUE benchmark.
Zero-Shot Dynamic Quantization for Transformer Inference (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for quantizing models require calibration or modification of parameters . run-time inference of such large models is costly due to large computational requirements .
Approach: They propose a run-time method for quantizing BERT-like models to 8-bit integers . they demonstrate that the method can be used on many NLP tasks without calibration steps .
Outcome: The proposed method reduces the accuracy loss associated with quantizing BERT-like models to 8-bit integers.
Fast and Accurate Neural Machine Translation with Translation Memory (2021.acl-long)

Copied to clipboard

Challenge: Existing knowledge demonstrates the superiority of TM-based neural machine translation only on TM specialized tasks .
Approach: They propose a translation memory-based approach to machine translation using a single bilingual sentence as its TM.
Outcome: The proposed approach surpasses baselines on two general tasks and improves on the TM-specialized translation tasks.
LLM-FP4: 4-Bit Floating-Point Quantized Transformers (2023.emnlp-main)

Copied to clipboard

Challenge: Existing quantization solutions are integer-based and struggle with bit widths below 8 bits.
Approach: They propose a method for quantizing weights and activations in large language models down to 4-bit floating-point values in a post-training manner.
Outcome: The proposed method outperforms existing methods on common sense zero-shot reasoning tasks by 12.7 points.
Dynamic Stashing Quantization for Efficient Transformer Training (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive performance on a range of Natural Language Processing (NLP) tasks.
Approach: They propose a dynamic quantization strategy that reduces the amount of memory operations and reduces arithmetic cost by 20.95 on two translation tasks and three classification tasks.
Outcome: The proposed model reduces the amount of arithmetic operations by 20.95 and the number of DRAM operations by 2.55 on two translation tasks and three classification tasks.
A Frustratingly Easy Post-Training Quantization Scheme for LLMs (2023.emnlp-main)

Copied to clipboard

Challenge: Efficient inference is crucial for hyper-scale AI models, including large language models, as their parameter count continues to increase for enhanced performance.
Approach: They propose a quantization scheme that fully utilizes the Transformer structure used in large language models to minimize the frequency of DRAM access while exploiting the parallelism of operations.
Outcome: The proposed method minimizes the frequency of DRAM access while exploiting the parallelism of operations through a dense matrix format.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations