Challenge: Existing work clusters channels using token dimension, which is suboptimal for grouping . a common challenge in LLM quantization is supporting "group-wise" quantization .
Approach: They propose a method to group channels with similar activation distributions using tokens . they propose shuffle operation that reduces relative GSM8K error by 86% .
Outcome: The proposed method reduces GSM8K error by 86% in both INT4 and MXFP4 formats compared to baselines .

Similar Papers

FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for quantization of large language models struggle to adapt to dynamic workloads.
Approach: a new framework optimizes the trade-off between inference speed and accuracy . FlexQuant enables fine-grained, layer-wise mixed-precision quantization .
Outcome: a new framework optimizes the trade-off between inference speed and accuracy . it achieves a 1.3 speedup across diverse language tasks with negligible accuracy loss .
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs (2026.acl-long)

Copied to clipboard

Challenge: NVFP4 supports fine-grained block isolation, 4-bit quantization errors and mixed-precision approaches . ARCQuant boosts NVFO4 performance via Augmented Residual Channels .
Approach: They propose a framework that boosts NVFP4 performance via Augmented Residual Channels.
Outcome: ARCQuant boosts NVFP4 performance via Augmented Residual Channels . the proposed framework achieves state-of-the-art accuracy comparable to full-precision baselines compared to FP16 .
MobileQuant: Mobile-friendly Quantization for On-device Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized language processing, but deployment on edge devices is costly in terms of memory, computation and energy.
Approach: They propose to reduce the number of bits used to represent weights and activations . they propose to use 8-bit activations to enable LLMs to fully exploit mobile-friendly hardware .
Outcome: The proposed method reduces the number of bits used to represent weights and activations . 8-bit activations are attractive for on-device deployment as they would exploit mobile-friendly hardware .
PreQuant: A Task-agnostic Quantization Approach for Pre-trained Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Quantization is a viable solution for pre-trained language models, but most existing methods are task-specific and require customized training and quantization with a large number of trainable parameters.
Approach: They propose a "quantize before fine-tuning" framework that allows for quantization with a large number of trainable parameters on each individual task.
Outcome: The proposed framework is compatible with quantization-aware training and post-training quantization and corrects quantization errors.
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for building efficient large language models with sub 2-bit weights are lacking in accuracy and scalability.
Approach: They propose a method that decouples parameters by splitting linear layers into two specialized branches.
Outcome: The proposed method achieves state-of-the-art performance in extremely low-bit quantization.
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in activation quantization methods cause outliers in tokens, causing extra overhead and speedup . a method to quantize per-tensor activation is currently challenging due to the outlier activation outlier.
Approach: They propose a method to find a set of key-value cache which mitigates outliers in subsequent tokens when inserted as a prefix.
Outcome: The proposed method surpasses the established baseline of per-tensor activation quantization and can be seamlessly integrated with the recent activation quantitative method.
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models (2026.eacl-short)

Copied to clipboard

Challenge: Large language models have limited truthfulness and tendency toward overconfidence constrain reliability in factual tasks.
Approach: They propose an efficient method that leverages semantic information encoded in LLMs to quantify uncertainty.
Outcome: The proposed method achieves comparable performance to baselines while significantly reducing computational overhead.
PE-QAT: Parameter-Efficient Quantization-Aware Training for Large Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Quantization Aware Training (QAT) is expensive to train and unscalable to large models.
Approach: They propose a parameter-efficient framework targeting per-channel 4-bit weight-activation quantization of large language models.
Outcome: The proposed framework preserves accuracy within 0.11 percentage points of the full-precision baseline on Llama-2-7B zero-shot tasks while training only 1.26% of total parameters.
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks. however, their extensive memory requirements present significant challenges for deployment in resource-constrained environments.
Approach: They propose a training-free framework that achieves ultra-low equivalent bit-width KV cache quantization.
Outcome: The proposed framework outperforms state-of-the-art methods on TruthfulQA and LongBench.
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression (2025.findings-emnlp)

Copied to clipboard

Challenge: storing more tokens in the KV cache at lower precision can enhance the long-context performance of large language models.
Approach: They propose a token-precision trade-off strategy to optimize KV cache compression . they also propose storing more tokens in the KV at lower precision .
Outcome: The proposed method achieves an optimal point within the Information Bottleneck compared to standalone KV pruning or KV quantization.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations