Robots Atlas>ROBOTS ATLAS
Inference

K-quants

2023ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Block quantization with 256-weight super-blocks and separately quantized scales and mins, giving a better quality/size trade-off than the earlier Q4_0/Q4_1 formats for local LLM inference.
Category
Inference
Abstraction level
Pattern
Operation level
InferenceServingDeployment
Use cases
Local LLM inference on CPURunning large models on GPUs with limited VRAMDistributing quantized models in the GGUF formatInference on edge devices and laptopsCompressing open-weight models (Llama, Mistral, Qwen, etc.)

How it works

1) A tensor's weights are split into 256-weight super-blocks. 2) A super-block is subdivided into smaller blocks: 16 blocks of 16 weights (Q2_K, Q3_K, Q6_K) or 8 blocks of 32 weights (Q4_K, Q5_K). 3) Each block has its own scale; "type-1" types (Q2_K, Q4_K, Q5_K) also have their own min (formula w = q · block_scale + block_min), while "type-0" types (Q3_K, Q6_K) use only a scale (w = q · block_scale). 4) The scales and mins are themselves quantized separately at low precision (e.g. 4-bit for Q2_K, 6-bit for Q3_K/Q4_K/Q5_K, 8-bit for Q6_K), yielding effective bits-per-weight of about 2.6 (Q2_K), 3.4375 (Q3_K), 4.5 (Q4_K), 5.5 (Q5_K), 6.5625 (Q6_K). 5) The _S/_M/_L variants assign different K-quant types to different model tensors, balancing size and quality. 6) At inference, blocks are dequantized on the fly for matrix multiplications; Q8_K is used to quantize intermediate results in dot products.

Problem solved

Earlier round-to-nearest formats (Q4_0, Q4_1) with a single scale per 32-weight block lost significant quality at low bits-per-weight. K-quants reduce this quality loss at comparable or smaller size, making it feasible to run large LLMs on consumer hardware (CPU, limited GPU memory).

Components

Super-block (256 weights)Grouping weights and shared quantization of scales/mins

The fundamental K-quant unit covering 256 weights, further subdivided into smaller blocks. Requires the tensor row size to be divisible by 256.

Inner block (16 or 32 weights)Local scale and min for a subset of weights

A smaller block inside the super-block: 16 weights (Q2_K, Q3_K, Q6_K) or 32 weights (Q4_K, Q5_K). Each block has its own scale (and min in type-1).

Quantized scales and minsReconstructing weight values: w = q · scale (+ min)

Scales (and mins in type-1 types) are quantized separately at low precision: 4-bit (Q2_K), 6-bit (Q3_K/Q4_K/Q5_K), 8-bit (Q6_K). They account for the overhead above the nominal quantization bit-width.

_S/_M/_L mixesPer-tensor precision selection

Predefined mixes (small/medium/large) assigning different K-quant types to different tensors (attention, feed-forward, output), e.g. Q4_K_M, Q5_K_M, Q3_K_L, to optimize the perplexity/size trade-off.

Official

Implementation

Implementation pitfalls
Divisibility-by-256 requirementMedium

The tensor row size must be divisible by 256 (the super-block size). Tensors that do not satisfy this are quantized with a different type (e.g. a Q8_0 fallback).

Strong quality drop at Q2_KHigh

The lowest types (especially Q2_K) raise perplexity significantly; without an importance matrix (imatrix), the quality of small models can drop noticeably.

Confusing _S/_M/_L labelsLow

The _S/_M/_L variants are per-tensor mixes of several types, not a single bit-width; the real size and quality depend on the specific mix, not just the leading number (e.g. Q4_K_M vs Q4_K_S).

Evolution

2023
K-quants introduced in llama.cpp (PR #1684)
Inflection point

Iwan Kawrakow adds the Q2_K–Q6_K and Q8_K types together with the _S/_M/_L mixes.

2024
Importance matrix (imatrix) integration

Importance-matrix-guided quantization improves K-quant (and I-quant) quality at low bits-per-weight.

2024
I-quants introduced as a complement

The new IQ family (IQ2_XXS, IQ3_S, IQ4_NL, etc.), based on the importance matrix, reaches lower bits-per-weight than K-quants.

Computational complexity

Space complexity: ~2.6–6.6 bit/wagę.

Parallelism

Parallelism level
Fully parallel

Block dequantization is independent across blocks/super-blocks, so it parallelizes easily on CPU (SIMD) and GPU.

Scope
Inference

Hardware requirements

Primary

K-quants were designed primarily for efficient CPU inference (AVX2/AVX-512, ARM NEON) in llama.cpp/GGML.

Good fit

Optimized CUDA/Metal dequantization kernels exist, enabling GPU inference and offloading.