Robots Atlas>ROBOTS ATLAS
Inference

NVFP4

2025ActivePublished
Key innovation
Introduces two-level scaling for the 4-bit FP4 format: a shared FP8 (E4M3) scale per small 16-element block plus a global FP32 per-tensor scale, yielding higher accuracy than MXFP4 at the same data bit-width.
Category
Inference
Abstraction level
Pattern
Operation level
ModelTrainingInference
Use cases
4-bit inference of large language models with accuracy close to FP8Low-precision pretraining of large language models on Blackwell hardwareReducing memory footprint and bandwidth versus 8-bit formatsServing reasoning models (e.g. DeepSeek-R1) with minimal quality loss

How it works

The tensor is first normalized by a single global FP32 factor across the whole tensor. It is then split into blocks of 16 consecutive elements; each block gets its own scaling factor in the FP8 E4M3 format (4-bit exponent, 3-bit mantissa — a fractional scale, not only powers of two). Elements are scaled and stored in the 4-bit FP4 E2M1 format, which encodes 16 values (0, ±0.5, ±1, ±1.5, ±2, ±3, ±4, ±6). Matrix multiplications run on fifth-generation NVIDIA Blackwell Tensor Cores, which natively multiply FP4 values, apply the block and tensor scales, and accumulate in higher precision. For training (NVIDIA's paper) the format is complemented by techniques such as random Hadamard transforms for outliers, two-dimensional quantization, stochastic rounding of gradients, and keeping sensitive layers in higher precision.

Problem solved

4-bit quantization with a single per-tensor factor, or with a coarse power-of-two scale (as in MXFP4), loses too much accuracy at such a low bit-width. NVFP4 reduces this loss via smaller blocks (16 elements) and a more precise, fractional FP8 E4M3 scale backed by a global FP32 normalization, delivering low memory footprint with accuracy close to FP8.

Components

Micro-scaling block (16 elements)Grouping unit for fine-grained block scaling.

A group of 16 consecutive tensor values that share a single FP8 E4M3 scaling factor.

Per-block scale (FP8 E4M3)Fine-grained fit of the block value range.

An 8-bit floating-point factor (4 exponent bits, 3 mantissa bits) that scales a 16-element block; a fractional scale, not only powers of two.

Global per-tensor scale (FP32)Global tensor-range normalization (the second scaling level).

A single FP32 factor for the whole tensor, normalizing its overall range before block scaling.

FP4 element (E2M1)Representation of individual elements after scaling.

A 4-bit floating-point value: 1 sign bit, 2 exponent bits, 1 mantissa bit; encodes 16 levels, maximum ±6.

Implementation

Implementation pitfalls
Two-level scaling overhead (16-block + FP32)Medium

NVFP4 uses a smaller block (16) with a per-block FP8 E4M3 scale plus a global FP32 scale. This gives higher accuracy than MXFP4 but a larger scale-bit overhead (~4.5 bits/element) and a more complex quantization/dequantization path.

Fix:Use the format on hardware with native FP4 support (Blackwell); weigh the overhead/accuracy trade-off against MXFP4 per model.
Outliers at such low precisionHigh

FP4 E2M1 encodes only 16 values, so outliers strongly degrade quality, especially in training. NVIDIA's paper applies random Hadamard transforms, 2D quantization, stochastic rounding of gradients, and keeping sensitive layers in higher precision.

Fix:Quantize selectively (mixed precision), suppress outliers (Hadamard), and use stochastic rounding during training.

Evolution

Original paper · 2025 · arXiv 2025 · NVIDIA
Pretraining Large Language Models with NVFP4
NVIDIA, Felix Abecassis, Anjulie Agrusa
2025
NVIDIA introduces NVFP4 for inference
Inflection point

NVIDIA presents NVFP4 as a 4-bit format with two-level scaling (FP8 E4M3 scale per 16-element block + a global FP32 scale) for efficient and accurate low-precision inference on the Blackwell architecture.

2025
LLM pretraining in NVFP4 (12B / 10T tokens)
Inflection point

NVIDIA's paper "Pretraining Large Language Models with NVFP4" demonstrates training a 12B model on 10T tokens in 4-bit precision, with quality comparable to an FP8 baseline, using random Hadamard transforms, 2D quantization, and stochastic rounding.

Hyperparameters (configurable axes)

Block sizeCritical

Number of elements sharing one scaling factor. In NVFP4 it is 16 (half of MXFP4).

16NVFP4 block size.
Element formatCritical

Single-element format: FP4 E2M1 (2 exponent bits, 1 mantissa bit).

E2M1Same element as MXFP4.
Block scale formatCritical

Format of the per-block scale; in NVFP4 it is FP8 E4M3 (fractional) instead of MXFP4's E8M0.

E4M3 (FP8)4 exponent bits, 3 mantissa bits.
Tensor scale formatHigh

Format of the global per-tensor factor; in NVFP4 it is FP32.

FP32The second (global) scaling level.

Computational complexity

Space complexity: ~4 bity/element + skala FP8 (E4M3) na 16 elementów + jeden FP32 na tensor.

Compute bottleneck

Memory bandwidth and dequantization

NVFP4 aims to relieve memory bandwidth and capacity via 4-bit weight storage. The bottleneck shifts toward Tensor Core arithmetic throughput and — on hardware without native FP4 support — toward the cost of dequantizing blocks and applying the two scaling levels.

Depends on
Natywne wsparcie sprzętowe (Blackwell, 5. gen. Tensor Cores)Obsługa dwupoziomowego skalowania (blok 16 + tensor FP32)

Execution paradigm

Primary mode
Dense

NVFP4 is a numeric format used in matrix multiplications on Tensor Cores; scaling operates at the 16-element block and whole-tensor level, and accumulation happens in higher precision.

Parallelism

Parallelism level
Fully parallel

Block scaling and NVFP4 matmuls are fully parallel on GPU Tensor Cores; the format is used for both inference and training.

Scope
InferenceTrainingAcross devices

Hardware requirements

Primary

Fifth-generation Tensor Cores in the NVIDIA Blackwell architecture natively support NVFP4 matmuls with two-level scaling (16-element block + tensor).

Possible

The format can be emulated in software, but without hardware acceleration or the performance benefits.

Limited

CPUs lack native FP4 matmuls; use is limited to emulation or dequantization to higher precision.