Robots Atlas>ROBOTS ATLAS
Inference

I-quants

2024ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Very-low-bit-per-weight quantization (from ~1.56 bpw) based on QuIP#-style lattice coding (E8 lattice, codebook) guided by an importance matrix, giving better quality than K-quants at the same size.
Category
Inference
Abstraction level
Pattern
Operation level
InferenceServingDeployment
Use cases
Extreme compression of large LLMs (1.5–3 bits per weight)Running very large models on GPUs with limited VRAMLocal LLM inference on laptops and edge devicesDistributing quantized models in the GGUF formatFitting larger models into the same memory with better quality than K-quants

How it works

1) A tensor's weights are split into 256-weight super-blocks with a shared scale (super_block_scale). 2) Instead of round-to-nearest quantization, groups of weights are encoded via a codebook based on the E8 lattice — a predefined set of points that weights are matched to; e.g. IQ2_XXS uses a 256-point codebook selected by occurrence frequency. 3) Signs are encoded separately (a sign-flipping strategy keeping an even count of negative signs per group), saving bits. 4) The quant selection is guided by an importance matrix (imatrix) collected on calibration text — weights that matter more for the model output are reproduced more accurately. 5) Effective bits-per-weight depend on the variant: ~1.56 (IQ1_S), 1.75 (IQ1_M), 2.06 (IQ2_XXS), 2.31 (IQ2_XS), 2.5 (IQ2_S), 3.06 (IQ3_XXS), 3.44 (IQ3_S), 4.25 (IQ4_XS); IQ4_NL is a 4-bit non-linear variant. 6) At inference, weights are reconstructed from codebook indices and the scale; codebook lookups are costlier than the simple multiply in K-quants, hence slower inference on some hardware.

Problem solved

K-quants lose quality at the lowest settings (especially 2–3 bits per weight), yet running very large models on consumer hardware demands even stronger compression. I-quants allow going down to ~1.5–2 bits per weight with less quality loss than similarly sized K-quants, making it possible to fit larger models into limited memory.

Components

Codebook (E8 lattice)Representing weight magnitudes at a very low bit count

A predefined set of points based on the highly symmetric E8 lattice (e.g. a 256-point codebook for IQ2_XXS), to which groups of weights are matched. The idea is borrowed from QuIP#.

Importance matrix (imatrix)Steering per-weight precision

Activation statistics collected on calibration text, indicating which weights most strongly affect the model output; they steer quant selection to minimize quality loss. Practically required for the lowest IQ types.

Official

Super-block (256 weights)Grouping weights and a shared scale

The fundamental quantization unit covering 256 weights with a shared scale (super_block_scale). Requires the tensor row size to be divisible by 256.

Sign encoding (sign-flipping)Saving bits on sign encoding

A sign-encoding strategy keeping an even count of negative elements per group, so one sign is derived from the others and bits are saved. Borrowed from QuIP#.

Official

Implementation

Implementation pitfalls
Slower inference on some hardwareMedium

Codebook lookups increase compute overhead; on CPUs and weaker GPUs I-quants can be slower than similarly sized K-quants.

Fix:Benchmark throughput on target hardware; if speed is the priority, consider K-quants.
Dependence on the importance matrixHigh

The lowest types (IQ1/IQ2) lose significant quality without a good importance matrix; results depend on the representativeness of the calibration data.

Fix:Generate the imatrix on representative text; avoid extremely low types without an imatrix.
Divisibility-by-256 requirementLow

The tensor row size must be divisible by 256; tensors that do not satisfy this are quantized with a different fallback type.

Fix:Rely on the converter's automatic fallback for unusual tensor shapes.

Evolution

2024
Theoretical basis: QuIP# (E8 lattice, codebook)

The QuIP# paper (Tseng et al., ICML 2024) formalizes LLM quantization using the E8 lattice and Hadamard incoherence; its ideas inspired I-quants.

2024
I-quants introduced in llama.cpp (PR #4773)
Inflection point

Iwan Kawrakow adds IQ2_XXS and IQ2_XS ("SOTA 2-bit quants") with QuIP#-style coding and an importance matrix.

2024
IQ4_NL — 4-bit non-linear variant (PR #5590)

The family is extended with 4-bit non-linear types (IQ4_NL, later IQ4_XS ~4.25 bpw).

2024
1-bit types: IQ1_S and IQ1_M (PR #6302)

Addition of the extremely low-bit types IQ1_S (~1.56 bpw) and IQ1_M (~1.75 bpw).

Hyperparameters (configurable axes)

Target bits-per-weight / IQ typeCritical

Choice of variant (IQ1_S … IQ4_XS) setting the effective bits-per-weight and the quality/size trade-off.

IQ2_XXS (~2.06 bpw)
IQ3_S (~3.44 bpw)
IQ4_XS (~4.25 bpw)
Importance matrix (imatrix) and calibration dataHigh

Presence and quality of the importance matrix and representativeness of the calibration text; practically required for the lowest IQ types.

z imatrix (zalecane)Significantly improves quality at low bits.
bez imatrixPossible for higher types, risky for IQ1/IQ2.

Computational complexity

Space complexity: ~1.56–4.25 bit/wagę.

Parallelism

Parallelism level
Fully parallel

Super-block decoding is independent, so it parallelizes easily; the bottleneck tends to be codebook lookups rather than inter-block dependencies.

Scope
Inference

Hardware requirements

Good fit

CUDA/Metal decode kernels exist; e.g. IQ2_XXS reaches about 155 t/s on an RTX-4080 (CUDA) and 54 t/s on an M2 Max (Metal) for Mistral-7B.

Possible

Codebook lookups are costlier than the simple multiply in K-quants, so I-quant inference can be slower on CPUs.