Abstract
Post-Training Quantization (PTQ) has emerged as an effective approach to reduce memory and computational demands during LLMs inference. However, existing PTQ methods are highly sensitive to ultra-low-bit quantization with significant performance loss, which is further exacerbated by recently released advanced models like LLaMA-3 and LLaMA-3.1. To address this challenge, we propose a novel PTQ framework, termed FRM-PTQ, by introducing feature relationship matching. This approach integrates token-level relationship modeling and structure-level distribution alignment based on the intra-block self-distillation framework to effectively mitigate significant performance degradation caused by low-bit quantization. Unlike conventional MSE loss methods, which focus solely on point-to-point discrepancies, feature relationship matching captures feature representations in high-dimensional spaces to effectively bridge the representation gap between quantized and full-precision blocks. Additionally, we propose a multi-granularity per-group quantization technique featuring a customized kernel, designed based on the quantization sensitivity of decoder block, to further relieve the quantization performance degradation. Extensive experimental results demonstrate that our method achieves outstanding performance in the W4A4 low-bit scenario, maintaining near full-precision accuracy while delivering a 2 × throughput improvement and a 3.17 × memory reduction. This advantage is particularly evident in the latest models such as LLaMA-3, LLaMA-3.1 and Qwen2.5 models, as well as in the W3A3 extreme low-bit scenarios. Codes are available at https://github.com/HITSZ-Miao-Group/FRM.
| Original language | English |
|---|---|
| Article number | 108619 |
| Journal | Neural Networks |
| Volume | 198 |
| DOIs | |
| State | Published - Jun 2026 |
| Externally published | Yes |
Keywords
- Large language model
- Post-Training quantization
Fingerprint
Dive into the research topics of 'FRM-PTQ: Feature relationship matching enhanced low-bit post-training quantization for large language models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver