K-quants
How it works
1) A tensor's weights are split into 256-weight super-blocks. 2) A super-block is subdivided into smaller blocks: 16 blocks of 16 weights (Q2_K, Q3_K, Q6_K) or 8 blocks of 32 weights (Q4_K, Q5_K). 3) Each block has its own scale; "type-1" types (Q2_K, Q4_K, Q5_K) also have their own min (formula w = q · block_scale + block_min), while "type-0" types (Q3_K, Q6_K) use only a scale (w = q · block_scale). 4) The scales and mins are themselves quantized separately at low precision (e.g. 4-bit for Q2_K, 6-bit for Q3_K/Q4_K/Q5_K, 8-bit for Q6_K), yielding effective bits-per-weight of about 2.6 (Q2_K), 3.4375 (Q3_K), 4.5 (Q4_K), 5.5 (Q5_K), 6.5625 (Q6_K). 5) The _S/_M/_L variants assign different K-quant types to different model tensors, balancing size and quality. 6) At inference, blocks are dequantized on the fly for matrix multiplications; Q8_K is used to quantize intermediate results in dot products.
Problem solved
Earlier round-to-nearest formats (Q4_0, Q4_1) with a single scale per 32-weight block lost significant quality at low bits-per-weight. K-quants reduce this quality loss at comparable or smaller size, making it feasible to run large LLMs on consumer hardware (CPU, limited GPU memory).
Components
The fundamental K-quant unit covering 256 weights, further subdivided into smaller blocks. Requires the tensor row size to be divisible by 256.
A smaller block inside the super-block: 16 weights (Q2_K, Q3_K, Q6_K) or 32 weights (Q4_K, Q5_K). Each block has its own scale (and min in type-1).
Scales (and mins in type-1 types) are quantized separately at low precision: 4-bit (Q2_K), 6-bit (Q3_K/Q4_K/Q5_K), 8-bit (Q6_K). They account for the overhead above the nominal quantization bit-width.
Predefined mixes (small/medium/large) assigning different K-quant types to different tensors (attention, feed-forward, output), e.g. Q4_K_M, Q5_K_M, Q3_K_L, to optimize the perplexity/size trade-off.
Official
Implementation
The tensor row size must be divisible by 256 (the super-block size). Tensors that do not satisfy this are quantized with a different type (e.g. a Q8_0 fallback).
The lowest types (especially Q2_K) raise perplexity significantly; without an importance matrix (imatrix), the quality of small models can drop noticeably.
The _S/_M/_L variants are per-tensor mixes of several types, not a single bit-width; the real size and quality depend on the specific mix, not just the leading number (e.g. Q4_K_M vs Q4_K_S).
Evolution
Iwan Kawrakow adds the Q2_K–Q6_K and Q8_K types together with the _S/_M/_L mixes.
Importance-matrix-guided quantization improves K-quant (and I-quant) quality at low bits-per-weight.
The new IQ family (IQ2_XXS, IQ3_S, IQ4_NL, etc.), based on the importance matrix, reaches lower bits-per-weight than K-quants.
Computational complexity
Space complexity: ~2.6–6.6 bit/wagę.
Parallelism
Block dequantization is independent across blocks/super-blocks, so it parallelizes easily on CPU (SIMD) and GPU.
Hardware requirements
K-quants were designed primarily for efficient CPU inference (AVX2/AVX-512, ARM NEON) in llama.cpp/GGML.
Optimized CUDA/Metal dequantization kernels exist, enabling GPU inference and offloading.