FP16
How it works
FP16 encodes a value as sign x 1.mantissa x 2^(exponent - 15), with a 5-bit exponent (bias 15) and a 10-bit mantissa plus an implicit leading bit (11 bits of precision). In mixed-precision training, weights, activations and gradients are stored in FP16, matrix multiplications run on Tensor Cores, and product accumulation is done in FP32; an FP32 master copy of the weights is also kept. To keep small gradients from underflowing FP16's narrow range, the loss is multiplied by a scaling factor before backpropagation and divided afterward (loss scaling). In inference, FP16 serves as a lightweight format for weights and activations that reduces VRAM usage.
Problem solved
FP32 training and inference consume large amounts of memory and bandwidth, and Tensor Cores reach peak throughput only in reduced precision. FP16 halves tensor size and speeds up matrix multiplication, but its narrow 5-bit exponent (max ~65504) causes overflow and gradient underflow, so FP16 training requires loss scaling and FP32 accumulation.
Key mechanisms
Strengths & limitations
Components
1 bit determining the sign of the value: 0 = positive, 1 = negative.
5 bits with a bias of 15 (exponent range -14 to +15), giving a maximum value of ~65504 and a narrow dynamic range versus FP32 and BF16.
10 explicitly stored bits plus 1 implicit leading bit (11 bits of precision) — more than BF16's 8 bits, so FP16 has higher relative precision at a comparable size.
Implementation
The narrow 5-bit exponent (max ~65504, min normal ~6e-5) causes large activations to overflow to inf and small gradients to underflow to zero.
Both formats are 16 bits, but FP16 has more mantissa and less exponent than BF16; code assuming BF16's range will fail in FP16 and vice versa.
Evolution
The half-precision format is formalized in the IEEE 754 standard as a storage type.
The first generation of Tensor Cores performs FP16 matrix multiplication with FP32 accumulation, dramatically accelerating training.
The NVIDIA and Baidu paper formalizes mixed-precision training with FP16, an FP32 master copy of weights, and loss scaling.
Hyperparameters (configurable axes)
A factor multiplying the loss before backpropagation to protect small gradients from underflow.
The precision of partial sums in matrix multiplication; usually FP32 for stability.
Computational complexity
Time complexity: O(1) na operację elementarną (stały koszt na wartość). Space complexity: 2 bajty / wartość.
On the NVIDIA A100, FP16 matrix multiplication with FP32 accumulation reaches 312 TFLOPS (same as BF16), versus 19.5 TFLOPS in FP32 without Tensor Cores. FP16 mixed-precision training typically preserves FP32-level accuracy when loss scaling is applied, while roughly halving time and memory usage.
Compute bottleneck
FP16 performance depends on the throughput of mixed-precision units and memory; the format itself imposes no compute cost beyond casts and FP32 accumulation.
Execution paradigm
All compute paths are active; FP16 only changes representation precision.
FP16 introduces no routing or conditional execution; it is a dense-compute format.
Parallelism
FP16 is a data format; element-wise ops and matrix multiplications are fully parallel on SIMD/Tensor Core units.
Hardware requirements
Natively supported by NVIDIA Tensor Cores since the Volta architecture (V100, 2017); FP16 matrix multiplication with FP32 accumulation reaches many times the throughput of FP32.
Also supported by AMD GPUs (RDNA/CDNA) and Apple Metal (half); widely available across consumer and datacenter accelerators.