MXFP4
How it works
A tensor is split into blocks of 32 consecutive elements. Each block gets one shared scaling factor in E8M0 format (an 8-bit power-of-two exponent), typically derived from the block's maximum absolute value. Elements are then scaled and stored in the 4-bit FP4 E2M1 format, which encodes 16 values (0, ±0.5, ±1, ±1.5, ±2, ±3, ±4, ±6). Matrix multiplications run on Tensor Cores that natively multiply FP4 values while accounting for the block scales, with accumulation in higher precision. Because of FP4's extremely low precision, quality depends heavily on the choice of scaling factor and rounding algorithm.
Problem solved
Classic 4-bit quantization with a single per-tensor factor loses too much accuracy at so few bits. MXFP4 gives each 32-element block its own scaling factor, drastically cutting memory and bandwidth usage (e.g. running an MoE model on a single GPU) at an acceptable quality loss.
Components
A group of 32 consecutive tensor values that share a single scaling factor.
8-bit exponent-only value representing the per-block scaling factor as a power of two.
A 4-bit floating-point value: 1 sign bit, 2 exponent bits, 1 mantissa bit; encodes 16 levels, maximum ±6.
Implementation
The shared scaling factor can only take power-of-two values. Deriving it from the block max-abs and rounding can waste part of the dynamic range and increase quantization error; quality depends heavily on the scale-selection and element-rounding strategy.
FP4 E2M1 encodes only 16 values, so quantizing all weights and activations usually degrades quality. In practice (e.g. gpt-oss) mainly the Mixture-of-Experts layer weights are quantized to MXFP4, while sensitive layers are kept at higher precision.
Native MXFP4 matmuls are only available on Blackwell Tensor Cores. On older hardware MXFP4 must be dequantized or emulated, which negates the bandwidth advantage and can slow inference down.
Evolution
The OCP consortium (AMD, Arm, Intel, Meta, Microsoft, NVIDIA, Qualcomm) publishes the MX standard, defining among others the 4-bit MXFP4 format (E2M1) with an E8M0 scale per 32-element block.
NVIDIA Blackwell introduces hardware support for microscaling (MXFP8, MXFP6, MXFP4) directly in the Tensor Cores.
OpenAI releases the open-weight gpt-oss-20b and gpt-oss-120b models with native MXFP4 quantization of the Mixture-of-Experts weights, letting 20b fit within 16GB and 120b on a single 80GB GPU.
Hyperparameters (configurable axes)
Number of elements sharing one scaling factor. Fixed at 32 in the MX standard.
The element format in MXFP4 is FP4 E2M1 (2 exponent bits, 1 mantissa bit).
Format of the shared scaling factor; in MX this is E8M0 (powers of two).
Computational complexity
Space complexity: ~4 bity/element + 8 bitów na 32 elementy (E8M0).
Compute bottleneck
The main goal of MXFP4 is to relieve memory bandwidth and capacity via 4-bit weight storage. The bottleneck shifts toward Tensor Core arithmetic throughput and — on hardware without native support — toward the cost of dequantizing blocks on every use.
Execution paradigm
MXFP4 is a numeric format used in matrix multiplications on Tensor Cores; scaling operates at block level and accumulation happens in higher precision.
Parallelism
Block scaling and MXFP4 matmuls are fully parallel on GPU Tensor Cores.
Hardware requirements
NVIDIA Blackwell Tensor Cores natively support MXFP4 matmuls with block scaling.
The format can be emulated in software (e.g. the microxcaling library), but without hardware acceleration.
CPUs lack native MXFP4 matmuls; use is limited to emulation or dequantization to higher precision.