INT4
How it works
INT4 maps continuous weight values onto 16 discrete integer levels. For each group of weights (e.g. 32, 64 or 128 elements), a scale and optionally a zero-point are computed so the group's range fits into [0,15] (asymmetric) or [-8,7] (symmetric). During inference, weights are dequantized on the fly to FP16/FP32 just before matrix multiplication, or dedicated kernels multiply directly on the packed values. Group-wise quantization limits error because each group has its own scale matched to the local weight distribution. Methods such as GPTQ and AWQ choose the quantization to minimize the layer's output error rather than just the weight error.
Problem solved
The FP16 weights of large language models occupy tens or hundreds of gigabytes, exceeding the memory of single accelerators and limiting inference throughput, which is memory-bandwidth bound. INT4 solves this by compressing weights to 4 bits, so the model fits in less memory and fewer bytes must be read per token.
Key mechanisms
Strengths & limitations
Components
Each weight is coded as one of 16 integer values (0-15 or -8..7), occupying 4 bits instead of 16 or 32.
A floating-point factor mapping the integer range onto the real weight range of the group; usually stored in FP16.
An integer offset in asymmetric quantization that lets non-zero-centered weight distributions be mapped; omitted in symmetric quantization.
Implementation
Large outliers in activations or weights make naive 4-bit quantization lose accuracy sharply, especially in larger models.
A large group (e.g. per-tensor) reduces scale memory overhead but increases quantization error; too small a group increases overhead and can slow kernels.
Evolution
GPTQ proves that 100B+ models can be quantized to INT4 with minimal quality loss using a post-training method.
AWQ (activation-aware quantization) and QLoRA (fine-tuning over 4-bit NF4 weights) make INT4 the default format for local inference and cheap fine-tuning.
The Q4 formats (K-quants) in llama.cpp make it possible to run LLaMA models on laptops and CPUs, popularizing 4-bit inference.
Hyperparameters (configurable axes)
Number of weights sharing one scale; a smaller group improves accuracy at the cost of overhead.
Symmetric (no zero-point) or asymmetric (with zero-point) value mapping.
The algorithm choosing the quantized values.
Computational complexity
Time complexity: O(1) na wagę przy dekwantyzacji + koszt mnożenia macierzy. Space complexity: 0,5 bajta / waga (+ narzut skal/zero-point).
Methods like GPTQ and AWQ can quantize 7-70B models to 4 bits with a small perplexity increase (often below 0.1-0.5 on WikiText-2 for larger models) versus FP16. The practical gain is a ~4x reduction in weight memory; e.g. a 70B model in FP16 (~140 GB) fits in INT4 in ~35-40 GB, enabling a single 48 GB card.
Compute bottleneck
In INT4 inference the bottleneck is reading packed weights from memory and dequantizing them; reducing bytes per token is what yields the speedup.
Execution paradigm
All layers run densely; only the weight representation changes.
INT4 is a weight compression scheme; it introduces no routing or conditional activation.
Parallelism
Dequantization and matrix multiplications are fully parallel; groups are quantized independently.
Hardware requirements
Dedicated kernels (e.g. Marlin, ExLlama, bitsandbytes) run matmul with INT4 weights and FP16 activations, exploiting GPU memory bandwidth.
llama.cpp implements efficient INT4 (Q4) kernels on CPU with SIMD instructions (AVX2/AVX-512, NEON), enabling LLM inference without a GPU.