Robots Atlas>ROBOTS ATLAS
Inference

FP16

2008ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
A 16-bit IEEE 754 half-precision floating-point format (1 sign bit, 5 exponent bits, 10 mantissa bits) that halves memory and bandwidth versus FP32, enabling mixed-precision training and inference on Tensor Cores.
Category
Inference
Abstraction level
Pattern
Operation level
ModelArchitecture blockTrainingInference
Use cases
Mixed-precision training of deep networks on NVIDIA GPUs (since Volta)Reduced-precision inference for lower VRAM usageAccelerating matrix multiplication on Tensor CoresIntermediate format in quantization pipelines (dequantization to FP16 before compute)Graphics and GPU compute requiring half precision

How it works

FP16 encodes a value as sign x 1.mantissa x 2^(exponent - 15), with a 5-bit exponent (bias 15) and a 10-bit mantissa plus an implicit leading bit (11 bits of precision). In mixed-precision training, weights, activations and gradients are stored in FP16, matrix multiplications run on Tensor Cores, and product accumulation is done in FP32; an FP32 master copy of the weights is also kept. To keep small gradients from underflowing FP16's narrow range, the loss is multiplied by a scaling factor before backpropagation and divided afterward (loss scaling). In inference, FP16 serves as a lightweight format for weights and activations that reduces VRAM usage.

Problem solved

FP32 training and inference consume large amounts of memory and bandwidth, and Tensor Cores reach peak throughput only in reduced precision. FP16 halves tensor size and speeds up matrix multiplication, but its narrow 5-bit exponent (max ~65504) causes overflow and gradient underflow, so FP16 training requires loss scaling and FP32 accumulation.

Key mechanisms

Floating-point representation: sign x 1.mantissa x 2^(exponent - 15)
Mixed-precision training with an FP32 master copy of weights
FP32 accumulation of matrix products on Tensor Cores
Loss scaling against gradient underflow
Casting/dequantization between FP16 and FP32 in compute pipelines

Strengths & limitations

Strengths
✓Higher relative precision (10 mantissa bits) than BF16 at the same 16-bit size
✓Half the memory footprint and bandwidth versus FP32
✓Native, broad hardware support (Tensor Cores since Volta, AMD, Apple, mobile)
✓Multi-fold speedup of matrix multiplication on mixed-precision units
✓Mature library support (cuDNN, PyTorch AMP, TensorRT)
Limitations
✗The narrow 5-bit exponent (max ~65504) invites overflow and gradient underflow
✗Training usually requires loss scaling, static or dynamic
✗Lower precision than FP32 harms summation of many small values without FP32 accumulation
✗Easy to confuse with BF16 despite the identical size (different bit layout)
✗Older CPUs/GPUs without native FP16 emulate it in software, losing the benefit

Components

Sign bitSign of the value

1 bit determining the sign of the value: 0 = positive, 1 = negative.

Exponent fieldEncodes the order of magnitude (scale) of the value

5 bits with a bias of 15 (exponent range -14 to +15), giving a maximum value of ~65504 and a narrow dynamic range versus FP32 and BF16.

Mantissa (significand)Encodes the significant digits (precision)

10 explicitly stored bits plus 1 implicit leading bit (11 bits of precision) — more than BF16's 8 bits, so FP16 has higher relative precision at a comparable size.

Implementation

Implementation pitfalls
Overflow and gradient underflowHigh

The narrow 5-bit exponent (max ~65504, min normal ~6e-5) causes large activations to overflow to inf and small gradients to underflow to zero.

Fix:Use loss scaling (static or dynamic) and accumulate in FP32; consider BF16 when dynamic range is critical.
Confusing FP16 with BF16Medium

Both formats are 16 bits, but FP16 has more mantissa and less exponent than BF16; code assuming BF16's range will fail in FP16 and vice versa.

Fix:Choose the format deliberately: FP16 for precision with loss scaling, BF16 for range without loss scaling.

Evolution

Original paper · 2008 · IEEE · IEEE Microprocessor Standards Committee
IEEE Standard for Floating-Point Arithmetic (IEEE 754-2008)
IEEE Microprocessor Standards Committee
2008
IEEE 754-2008 standardizes binary16 (half precision)
Inflection point

The half-precision format is formalized in the IEEE 754 standard as a storage type.

2017
NVIDIA Volta (V100) introduces Tensor Cores for FP16
Inflection point

The first generation of Tensor Cores performs FP16 matrix multiplication with FP32 accumulation, dramatically accelerating training.

2018
Publication of Mixed Precision Training

The NVIDIA and Baidu paper formalizes mixed-precision training with FP16, an FP32 master copy of weights, and loss scaling.

Hyperparameters (configurable axes)

Loss scalingHigh

A factor multiplying the loss before backpropagation to protect small gradients from underflow.

static (np. 1024)A fixed factor.
dynamicAutomatically tuned during training.
Accumulation precisionHigh

The precision of partial sums in matrix multiplication; usually FP32 for stability.

FP32Default, protects accuracy.
FP16Faster but numerically risky.

Computational complexity

Computational characteristics
→2 bytes per value (half of FP32)
→Normalized range ~6e-5 to 65504
→11 bits of effective precision (10 + implicit bit)
→Tensor Core throughput typically 2x versus FP32 (FP32 accumulation)
→~50% reduction in memory traffic versus FP32

Time complexity: O(1) na operację elementarną (stały koszt na wartość). Space complexity: 2 bajty / wartość.

Benchmark notes

On the NVIDIA A100, FP16 matrix multiplication with FP32 accumulation reaches 312 TFLOPS (same as BF16), versus 19.5 TFLOPS in FP32 without Tensor Cores. FP16 mixed-precision training typically preserves FP32-level accuracy when loss scaling is applied, while roughly halving time and memory usage.

Compute bottleneck

Tensor Core and memory bandwidth

FP16 performance depends on the throughput of mixed-precision units and memory; the format itself imposes no compute cost beyond casts and FP32 accumulation.

Depends on
Wsparcie sprzętowe FP16 (Tensor Core)Precyzja akumulacji

Execution paradigm

Primary mode
Dense

All compute paths are active; FP16 only changes representation precision.

Activation pattern
All paths active
Routing mechanism

FP16 introduces no routing or conditional execution; it is a dense-compute format.

Parallelism

Parallelism level
Fully parallel

FP16 is a data format; element-wise ops and matrix multiplications are fully parallel on SIMD/Tensor Core units.

Scope
TrainingInference

Hardware requirements

Primary

Natively supported by NVIDIA Tensor Cores since the Volta architecture (V100, 2017); FP16 matrix multiplication with FP32 accumulation reaches many times the throughput of FP32.

Good fit

Also supported by AMD GPUs (RDNA/CDNA) and Apple Metal (half); widely available across consumer and datacenter accelerators.