Layer offloading
How it works
The model is split into layers (transformer blocks). Some layers stay permanently in VRAM while the rest are held in CPU RAM or on an NVMe disk. During the forward pass, when a layer outside the GPU is needed, its weights are copied over PCIe into VRAM, the computation runs, and the space is freed for the next layer; activations flow through successive layers sequentially. To hide transfer latency, prefetching is used (asynchronously loading the next layer while computing the current one) and transfer is overlapped with compute. Not only layer weights but also the KV cache and optimizer states (in training, e.g. ZeRO-Offload) can be offloaded. The number of layers kept on the GPU (e.g. the n-gpu-layers parameter in llama.cpp) controls the tradeoff between VRAM usage and speed.
Problem solved
Large models do not fit entirely in a single GPU's VRAM, and buying more accelerators is often impossible or uneconomical. Layer offloading solves this by keeping only part of the model in VRAM and storing the rest in cheaper, more abundant CPU memory or on an NVMe disk; layers are fetched to the GPU just before use. This enables inference (and partly training) of models larger than VRAM, at the cost of a speed drop caused by PCIe transfer.
Key mechanisms
Strengths & limitations
Components
Deciding which layers (or tensors) to keep in VRAM and which in CPU RAM or on disk, driven by the memory budget.
Copying a layer's weights from CPU memory/disk into VRAM just before compute and freeing them afterward.
Asynchronously loading the next layer while computing the current one, hiding transfer latency behind compute.
Implementation
Transferring weights over PCIe is many times slower than reading from VRAM; excessive offload dramatically lowers tokens per second.
With a long context, the KV cache can exceed VRAM even when weights are offloaded, causing out-of-memory errors.
Evolution
Microsoft DeepSpeed introduces offloading optimizer states and parameters to CPU/NVMe, enabling training of models many times larger than VRAM.
llama.cpp (n-gpu-layers), Hugging Face Accelerate (device_map) and FlexGen popularize layer offloading for inference on consumer hardware.
Hyperparameters (configurable axes)
How many layers to keep in VRAM; the rest are offloaded to CPU/disk.
Where offloaded data is stored: CPU memory or NVMe disk.
Computational complexity
Time complexity: t ~ t_compute + (bajty offloadowanych warstw) / przepustowosc PCIe. Space complexity: VRAM ~ (warstwy na GPU / wszystkie warstwy) x rozmiar modelu.
PCIe 4.0 x16 bandwidth is ~32 GB/s, versus ~1-3 TB/s VRAM bandwidth of modern GPUs - so each offloaded layer significantly slows generation. In practice llama.cpp lets you smoothly tune n-gpu-layers: the more layers fit in VRAM, the higher the tokens/s; full CPU offload reduces speed to CPU-inference levels.
Compute bottleneck
The bottleneck is the PCIe bus bandwidth between CPU/disk and GPU; it, not compute, limits speed under offload.
Execution paradigm
All layers execute; the difference is the memory location of the weights.
Offloading does not change the compute graph or introduce routing; it only controls where weights are stored and transferred.
Parallelism
Transfer and compute can overlap (prefetching), but inter-layer dependency forces sequential activation flow; cross-device parallelism is conditional on the bus.
Hardware requirements
The GPU runs the layer compute, but the benefit of offload depends on a fast bus (PCIe 4.0/5.0, NVLink) and large, fast CPU memory.
Offloaded layers can be computed directly on the CPU (e.g. hybrid mode in llama.cpp) when transferring to the GPU is not worthwhile for a given layer.