I-quants
How it works
1) A tensor's weights are split into 256-weight super-blocks with a shared scale (super_block_scale). 2) Instead of round-to-nearest quantization, groups of weights are encoded via a codebook based on the E8 lattice — a predefined set of points that weights are matched to; e.g. IQ2_XXS uses a 256-point codebook selected by occurrence frequency. 3) Signs are encoded separately (a sign-flipping strategy keeping an even count of negative signs per group), saving bits. 4) The quant selection is guided by an importance matrix (imatrix) collected on calibration text — weights that matter more for the model output are reproduced more accurately. 5) Effective bits-per-weight depend on the variant: ~1.56 (IQ1_S), 1.75 (IQ1_M), 2.06 (IQ2_XXS), 2.31 (IQ2_XS), 2.5 (IQ2_S), 3.06 (IQ3_XXS), 3.44 (IQ3_S), 4.25 (IQ4_XS); IQ4_NL is a 4-bit non-linear variant. 6) At inference, weights are reconstructed from codebook indices and the scale; codebook lookups are costlier than the simple multiply in K-quants, hence slower inference on some hardware.
Problem solved
K-quants lose quality at the lowest settings (especially 2–3 bits per weight), yet running very large models on consumer hardware demands even stronger compression. I-quants allow going down to ~1.5–2 bits per weight with less quality loss than similarly sized K-quants, making it possible to fit larger models into limited memory.
Components
A predefined set of points based on the highly symmetric E8 lattice (e.g. a 256-point codebook for IQ2_XXS), to which groups of weights are matched. The idea is borrowed from QuIP#.
Activation statistics collected on calibration text, indicating which weights most strongly affect the model output; they steer quant selection to minimize quality loss. Practically required for the lowest IQ types.
Official
The fundamental quantization unit covering 256 weights with a shared scale (super_block_scale). Requires the tensor row size to be divisible by 256.
A sign-encoding strategy keeping an even count of negative elements per group, so one sign is derived from the others and bits are saved. Borrowed from QuIP#.
Official
Implementation
Codebook lookups increase compute overhead; on CPUs and weaker GPUs I-quants can be slower than similarly sized K-quants.
The lowest types (IQ1/IQ2) lose significant quality without a good importance matrix; results depend on the representativeness of the calibration data.
The tensor row size must be divisible by 256; tensors that do not satisfy this are quantized with a different fallback type.
Evolution
The QuIP# paper (Tseng et al., ICML 2024) formalizes LLM quantization using the E8 lattice and Hadamard incoherence; its ideas inspired I-quants.
Iwan Kawrakow adds IQ2_XXS and IQ2_XS ("SOTA 2-bit quants") with QuIP#-style coding and an importance matrix.
The family is extended with 4-bit non-linear types (IQ4_NL, later IQ4_XS ~4.25 bpw).
Addition of the extremely low-bit types IQ1_S (~1.56 bpw) and IQ1_M (~1.75 bpw).
Hyperparameters (configurable axes)
Choice of variant (IQ1_S … IQ4_XS) setting the effective bits-per-weight and the quality/size trade-off.
Presence and quality of the importance matrix and representativeness of the calibration text; practically required for the lowest IQ types.
Computational complexity
Space complexity: ~1.56–4.25 bit/wagę.
Parallelism
Super-block decoding is independent, so it parallelizes easily; the bottleneck tends to be codebook lookups rather than inter-block dependencies.
Hardware requirements
CUDA/Metal decode kernels exist; e.g. IQ2_XXS reaches about 155 t/s on an RTX-4080 (CUDA) and 54 t/s on an M2 Max (Metal) for Mistral-7B.
Codebook lookups are costlier than the simple multiply in K-quants, so I-quant inference can be slower on CPUs.