Publication

Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression

返回学术发表
Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression 论文配图

Peiyu Liu, Ze-Feng Gao, Wayne Xin Zhao#, Yipeng Ma, Tao Wang, Ji-Rong Wen

Annual Meeting of the Association for Computational Linguistics (ACL2024), 2024

摘要

In this paper, we introduce DecoQuant, a novel data-free low-bit quantization technique based on tensor decomposition methods, to effectively compress KV cache. Our core idea is to adjust the outlier distribution of the original matrix by performing tensor decomposition, so that the quantization difficulties are migrated from the matrix to decomposed local tensors. Specially, we find that outliers mainly concentrate on small local tensors, while large tensors tend to have a narrower value range.