Robots Atlas>ROBOTS ATLAS
Inference

MXFP4

2023ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Makes 4-bit floating-point precision practical via a shared power-of-two exponent (E8M0) over blocks of 32 elements, capturing local magnitude variation and enabling large models to run at about 4.25 bits per parameter.
Category
Inference
Abstraction level
Pattern
Operation level
ModelTrainingInference
Use cases
Low-memory inference of large language modelsQuantizing Mixture-of-Experts weights (e.g. gpt-oss)Running models on consumer and edge hardwareReducing memory bandwidth and serving cost

How it works

A tensor is split into blocks of 32 consecutive elements. Each block gets one shared scaling factor in E8M0 format (an 8-bit power-of-two exponent), typically derived from the block's maximum absolute value. Elements are then scaled and stored in the 4-bit FP4 E2M1 format, which encodes 16 values (0, ±0.5, ±1, ±1.5, ±2, ±3, ±4, ±6). Matrix multiplications run on Tensor Cores that natively multiply FP4 values while accounting for the block scales, with accumulation in higher precision. Because of FP4's extremely low precision, quality depends heavily on the choice of scaling factor and rounding algorithm.

Problem solved

Classic 4-bit quantization with a single per-tensor factor loses too much accuracy at so few bits. MXFP4 gives each 32-element block its own scaling factor, drastically cutting memory and bandwidth usage (e.g. running an MoE model on a single GPU) at an acceptable quality loss.

Components

Microscaling block (32 elements)Grouping unit for shared scaling.

A group of 32 consecutive tensor values that share a single scaling factor.

Shared scale exponent (E8M0)Scales the block value range at low hardware cost.

8-bit exponent-only value representing the per-block scaling factor as a power of two.

FP4 element (E2M1)Representation of individual elements after scaling.

A 4-bit floating-point value: 1 sign bit, 2 exponent bits, 1 mantissa bit; encodes 16 levels, maximum ±6.

Implementation

Implementation pitfalls
E8M0 scale selection (powers of two only)High

The shared scaling factor can only take power-of-two values. Deriving it from the block max-abs and rounding can waste part of the dynamic range and increase quantization error; quality depends heavily on the scale-selection and element-rounding strategy.

Fix:Use careful scale calibration (e.g. searching for the optimal power of two) and advanced rounding (stochastic rounding / round-to-nearest-even).
Extremely low FP4 precision (16 levels, max ±6)High

FP4 E2M1 encodes only 16 values, so quantizing all weights and activations usually degrades quality. In practice (e.g. gpt-oss) mainly the Mixture-of-Experts layer weights are quantized to MXFP4, while sensitive layers are kept at higher precision.

Fix:Quantize selectively (mixed precision): only robust layers (e.g. MoE), keeping embeddings, norms and sensitive layers in BF16/FP16/FP8.
No native support outside NVIDIA BlackwellMedium

Native MXFP4 matmuls are only available on Blackwell Tensor Cores. On older hardware MXFP4 must be dequantized or emulated, which negates the bandwidth advantage and can slow inference down.

Fix:On hardware without native support, dequantize weights to BF16/FP16 at load time or use dedicated format-aware kernels (e.g. Triton).

Evolution

Original paper · 2023 · arXiv 2023 (OCP Microscaling Formats) · Bita Darvish Rouhani
Microscaling Data Formats for Deep Learning
Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall
2023
OCP MX v1.0 specification defines MXFP4
Inflection point

The OCP consortium (AMD, Arm, Intel, Meta, Microsoft, NVIDIA, Qualcomm) publishes the MX standard, defining among others the 4-bit MXFP4 format (E2M1) with an E8M0 scale per 32-element block.

2024
Native MXFP4 support in NVIDIA Blackwell Tensor Cores
Inflection point

NVIDIA Blackwell introduces hardware support for microscaling (MXFP8, MXFP6, MXFP4) directly in the Tensor Cores.

2025
OpenAI releases gpt-oss with MoE weights in MXFP4
Inflection point

OpenAI releases the open-weight gpt-oss-20b and gpt-oss-120b models with native MXFP4 quantization of the Mixture-of-Experts weights, letting 20b fit within 16GB and 120b on a single 80GB GPU.

Hyperparameters (configurable axes)

Block sizeCritical

Number of elements sharing one scaling factor. Fixed at 32 in the MX standard.

32Value defined in OCP MX v1.0.
Element formatCritical

The element format in MXFP4 is FP4 E2M1 (2 exponent bits, 1 mantissa bit).

E2M1The only FP4 element variant in the MX standard.
Scale formatHigh

Format of the shared scaling factor; in MX this is E8M0 (powers of two).

E8M08-bit exponent, no mantissa.

Computational complexity

Space complexity: ~4 bity/element + 8 bitów na 32 elementy (E8M0).

Compute bottleneck

Memory bandwidth and dequantization

The main goal of MXFP4 is to relieve memory bandwidth and capacity via 4-bit weight storage. The bottleneck shifts toward Tensor Core arithmetic throughput and — on hardware without native support — toward the cost of dequantizing blocks on every use.

Depends on
Natywne wsparcie sprzętowe (Blackwell)Obsługa skalowania blokowego

Execution paradigm

Primary mode
Dense

MXFP4 is a numeric format used in matrix multiplications on Tensor Cores; scaling operates at block level and accumulation happens in higher precision.

Parallelism

Parallelism level
Fully parallel

Block scaling and MXFP4 matmuls are fully parallel on GPU Tensor Cores.

Scope
InferenceAcross devices

Hardware requirements

Primary

NVIDIA Blackwell Tensor Cores natively support MXFP4 matmuls with block scaling.

Possible

The format can be emulated in software (e.g. the microxcaling library), but without hardware acceleration.

Limited

CPUs lack native MXFP4 matmuls; use is limited to emulation or dequantization to higher precision.