Robots Atlas>ROBOTS ATLAS
Inference

INT4

2022ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Representing model weights in 4-bit integers (16 levels), shrinking weights ~8x versus FP32 and ~4x versus FP16, making it possible to run large LLMs on a single GPU or CPU.
Category
Inference
Abstraction level
Pattern
Operation level
ModelArchitecture blockPost-trainingInference
Use cases
Running large LLMs on a single consumer GPULLM inference on CPUs and laptops (e.g. via llama.cpp)Reducing cost and VRAM usage in model servingDeploying models on edge devices with limited memoryIncreasing throughput of memory-bandwidth-bound inference

How it works

INT4 maps continuous weight values onto 16 discrete integer levels. For each group of weights (e.g. 32, 64 or 128 elements), a scale and optionally a zero-point are computed so the group's range fits into [0,15] (asymmetric) or [-8,7] (symmetric). During inference, weights are dequantized on the fly to FP16/FP32 just before matrix multiplication, or dedicated kernels multiply directly on the packed values. Group-wise quantization limits error because each group has its own scale matched to the local weight distribution. Methods such as GPTQ and AWQ choose the quantization to minimize the layer's output error rather than just the weight error.

Problem solved

The FP16 weights of large language models occupy tens or hundreds of gigabytes, exceeding the memory of single accelerators and limiting inference throughput, which is memory-bandwidth bound. INT4 solves this by compressing weights to 4 bits, so the model fits in less memory and fewer bytes must be read per token.

Key mechanisms

Mapping weights to 16 integer levels (0-15 or -8..7)
Group-wise quantization with a per-group scale
Zero-point in the asymmetric variant
On-the-fly dequantization to FP16 or kernels operating on packed weights
Format variants: NF4, K-quants, I-quants

Strengths & limitations

Strengths
✓~8x weight size reduction versus FP32 and ~4x versus FP16
✓Enables inference of large LLMs on a single GPU or CPU
✓Increases throughput of memory-bound inference (fewer bytes per token)
✓Rich ecosystem of methods (GPTQ, AWQ, NF4) and kernels (Marlin, ExLlama)
✓Lower serving and edge-deployment costs
Limitations
✗Naive 4-bit quantization sharply lowers quality without advanced methods
✗Activation outliers damage accuracy
✗Scale and zero-point overhead raises effective bits above 4
✗Requires dedicated kernels to realize the actual speedup
✗Lower accuracy than INT8/FP16, especially for small models

Components

4-bit integer levelCompressed weight representation

Each weight is coded as one of 16 integer values (0-15 or -8..7), occupying 4 bits instead of 16 or 32.

Per-group scaleDequantizing values back to the real scale

A floating-point factor mapping the integer range onto the real weight range of the group; usually stored in FP16.

Zero-pointCorrecting the offset of the weight distribution

An integer offset in asymmetric quantization that lets non-zero-centered weight distributions be mapped; omitted in symmetric quantization.

Implementation

Implementation pitfalls
Activation outliersHigh

Large outliers in activations or weights make naive 4-bit quantization lose accuracy sharply, especially in larger models.

Fix:Use activation-aware methods (AWQ), group-wise quantization with small groups, and mixed precision for outlier channels.
Group size too largeMedium

A large group (e.g. per-tensor) reduces scale memory overhead but increases quantization error; too small a group increases overhead and can slow kernels.

Fix:Tune the group size (typically 32-128) to balance accuracy, memory and kernel performance.

Evolution

Original paper · 2022 · ICLR 2023 · Elias Frantar
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh
2022
GPTQ demonstrates accurate LLM quantization to 4 and 3 bits
Inflection point

GPTQ proves that 100B+ models can be quantized to INT4 with minimal quality loss using a post-training method.

2023
AWQ and QLoRA popularize INT4 (NF4)
Inflection point

AWQ (activation-aware quantization) and QLoRA (fine-tuning over 4-bit NF4 weights) make INT4 the default format for local inference and cheap fine-tuning.

2023
llama.cpp popularizes Q4 quantization on CPU

The Q4 formats (K-quants) in llama.cpp make it possible to run LLaMA models on laptops and CPUs, popularizing 4-bit inference.

Hyperparameters (configurable axes)

Group sizeHigh

Number of weights sharing one scale; a smaller group improves accuracy at the cost of overhead.

128Typical accuracy/overhead balance.
32Higher accuracy, more metadata.
Quantization symmetryMedium

Symmetric (no zero-point) or asymmetric (with zero-point) value mapping.

asymmetricBetter for non-zero-centered distributions.
symmetricSimpler, lower overhead.
Quantization methodCritical

The algorithm choosing the quantized values.

GPTQHessian-based, minimizes output error.
AWQActivation-aware, protects salient channels.
NF4The normal-float format from QLoRA.

Computational complexity

Computational characteristics
→~4.3-4.5 bits per weight effectively (4 bits + metadata)
→16 discrete representation levels per value
→Group size typically 32-128 elements
→~4x reduction in per-token memory traffic versus FP16
→Requires INT4 kernels (Marlin, ExLlama, llama.cpp) for speedup

Time complexity: O(1) na wagę przy dekwantyzacji + koszt mnożenia macierzy. Space complexity: 0,5 bajta / waga (+ narzut skal/zero-point).

Benchmark notes

Methods like GPTQ and AWQ can quantize 7-70B models to 4 bits with a small perplexity increase (often below 0.1-0.5 on WikiText-2 for larger models) versus FP16. The practical gain is a ~4x reduction in weight memory; e.g. a 70B model in FP16 (~140 GB) fits in INT4 in ~35-40 GB, enabling a single 48 GB card.

Compute bottleneck

Memory reads and dequantization

In INT4 inference the bottleneck is reading packed weights from memory and dequantizing them; reducing bytes per token is what yields the speedup.

Depends on
Rozmiar grupy i narzut skalJakość kerneli INT4

Execution paradigm

Primary mode
Dense

All layers run densely; only the weight representation changes.

Activation pattern
All paths active
Routing mechanism

INT4 is a weight compression scheme; it introduces no routing or conditional activation.

Parallelism

Parallelism level
Fully parallel

Dequantization and matrix multiplications are fully parallel; groups are quantized independently.

Scope
InferenceAcross tokens

Hardware requirements

Primary

Dedicated kernels (e.g. Marlin, ExLlama, bitsandbytes) run matmul with INT4 weights and FP16 activations, exploiting GPU memory bandwidth.

Good fit

llama.cpp implements efficient INT4 (Q4) kernels on CPU with SIMD instructions (AVX2/AVX-512, NEON), enabling LLM inference without a GPU.