Papers by Zhongxue Gan
Boost Transformer-based Language Models with GPU-Friendly Sparsity and Quantization (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have shown that transformer-based language models can achieve state-of-the-art performance on various natural language processing tasks. |
| Approach: | They propose a GPU-friendly 2:4 fine-grained structured sparsity and quantization scheme to maximize the GPU's acceleration constraints. |
| Outcome: | The proposed scheme achieves state-of-the-art compression on TLM model of various encoder and decoder blocks with negligible accuracy degradation on SQuAD, GLUE, CNN-DM & XSum and WikiText benchmarking tasks. |