To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration

Zeyu Yang1 Tianyi Zhang1 Jianwen Xie2 Chuan Li2 Zhaozhuo Xu3 Anshumali Shrivastava1
1
2
3
ICLR 2026

TL;DR

GenAI model weights exhibit exponent concentrationFloating-point exponents in trained model weights cluster within a narrow range with entropy around only 2–3 bits, far below the 4–8 bits allocated by standard formats., enabling lossless FP8 compressionECF8 applies Huffman coding to low-entropy exponents with GPU-optimized decoding, achieving bit-exact weight reconstruction with zero quality degradation. that delivers real inference speedupUp to 26.9% memory savings free GPU memory for larger batch sizes, yielding up to 177.1% throughput acceleration on models with up to 671B parameters..

We present a theoretical and empirical study of exponent concentration in GenAI model weights: floating-point exponents consistently exhibit low entropy across architectures and modalities. We trace this phenomenon to $\alpha$-stable distributions induced by SGD and prove a compression limit near FP4.67. Building on these insights, we propose ECF8, a lossless FP8 compression framework with entropy-aware encoding and GPU-optimized decoding. Experiments on LLMs and DiTs with up to 671B parameters demonstrate up to 26.9% memory savings and 177.1% throughput acceleration, all with perfectly lossless computation.

26.9%
Memory Savings
177.1%
Throughput Acceleration
671B
Largest Model Tested
0
Output Deviation