Robots Atlas>ROBOTS ATLAS
Inference

Layer offloading

2021ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Distributing a model's layers (or their fragments and the KV cache) between GPU memory and slower but larger CPU memory (RAM) or NVMe disk, allowing a model larger than VRAM to run at the cost of data transfer over the PCIe bus.
Category
Inference
Abstraction level
Pattern
Operation level
InferenceServingDeployment
Use cases
Running models larger than VRAM on a single GPULLM inference on laptops and low-memory cards (llama.cpp n-gpu-layers)Hybrid GPU+CPU inference for MoE modelsTraining/fine-tuning with optimizer-state offload (ZeRO-Offload, ZeRO-Infinity)Cutting cost by using cheaper CPU/NVMe memory instead of extra GPUs

How it works

The model is split into layers (transformer blocks). Some layers stay permanently in VRAM while the rest are held in CPU RAM or on an NVMe disk. During the forward pass, when a layer outside the GPU is needed, its weights are copied over PCIe into VRAM, the computation runs, and the space is freed for the next layer; activations flow through successive layers sequentially. To hide transfer latency, prefetching is used (asynchronously loading the next layer while computing the current one) and transfer is overlapped with compute. Not only layer weights but also the KV cache and optimizer states (in training, e.g. ZeRO-Offload) can be offloaded. The number of layers kept on the GPU (e.g. the n-gpu-layers parameter in llama.cpp) controls the tradeoff between VRAM usage and speed.

Problem solved

Large models do not fit entirely in a single GPU's VRAM, and buying more accelerators is often impossible or uneconomical. Layer offloading solves this by keeping only part of the model in VRAM and storing the rest in cheaper, more abundant CPU memory or on an NVMe disk; layers are fetched to the GPU just before use. This enables inference (and partly training) of models larger than VRAM, at the cost of a speed drop caused by PCIe transfer.

Key mechanisms

Splitting layers between VRAM and CPU RAM / NVMe disk
Transferring weights over PCIe just before compute
Prefetching the next layer while computing the current one
Overlapping transfer with compute
Offloading the KV cache and optimizer states (training)

Strengths & limitations

Strengths
โœ“Enables running a model larger than VRAM
โœ“Uses cheap, abundant CPU/NVMe memory instead of extra GPUs
โœ“Tunable VRAM-vs-speed tradeoff (e.g. n-gpu-layers)
โœ“Prefetching and transfer overlap mitigate the overhead
โœ“Supports both inference and training (ZeRO-Offload/Infinity)
Limitations
โœ—PCIe transfer is many times slower than reading from VRAM
โœ—Excessive offload drastically lowers tokens/s
โœ—The KV cache at long context can still exceed VRAM
โœ—The gain depends on a fast bus (PCIe 4.0/5.0, NVLink)
โœ—Complexity of scheduling transfer and compute

Components

Layer placement across devicesSplitting the model across fast and slow memory

Deciding which layers (or tensors) to keep in VRAM and which in CPU RAM or on disk, driven by the memory budget.

PCIe transferDelivering data to the compute unit

Copying a layer's weights from CPU memory/disk into VRAM just before compute and freeing them afterward.

Prefetching and transfer overlapLimiting the performance penalty

Asynchronously loading the next layer while computing the current one, hiding transfer latency behind compute.

Implementation

Implementation pitfalls
PCIe bottleneckHigh

Transferring weights over PCIe is many times slower than reading from VRAM; excessive offload dramatically lowers tokens per second.

Fix:Keep as many layers on the GPU as VRAM allows; use prefetching and overlap transfer with compute.
Overlooking KV cache offloadMedium

With a long context, the KV cache can exceed VRAM even when weights are offloaded, causing out-of-memory errors.

Fix:Offload or quantize the KV cache, limit context length, or use grouped attention (GQA/MQA).

Evolution

Original paper ยท 2021 ยท USENIX ATC 2021 ยท Jie Ren
ZeRO-Offload: Democratizing Billion-Scale Model Training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Yuxiong He
2021
ZeRO-Offload and ZeRO-Infinity
Inflection point

Microsoft DeepSpeed introduces offloading optimizer states and parameters to CPU/NVMe, enabling training of models many times larger than VRAM.

2023
Layer offloading in local LLM inference

llama.cpp (n-gpu-layers), Hugging Face Accelerate (device_map) and FlexGen popularize layer offloading for inference on consumer hardware.

Hyperparameters (configurable axes)

GPU layer countCritical

How many layers to keep in VRAM; the rest are offloaded to CPU/disk.

0Full CPU inference.
35Some layers on GPU (hybrid).
allWhole model in VRAM, no offload.
Offload targetHigh

Where offloaded data is stored: CPU memory or NVMe disk.

CPU RAMFaster than disk, limited by RAM capacity.
NVMeLargest capacity, slowest transfer.

Computational complexity

Computational characteristics
โ†’VRAM usage proportional to the number of layers on GPU
โ†’Constraint: total CPU/NVMe memory, not just VRAM
โ†’Speed strongly dependent on PCIe bandwidth
โ†’Prefetching reduces visible transfer latency
โ†’Hybrid mode: some layers computed on the CPU

Time complexity: t ~ t_compute + (bajty offloadowanych warstw) / przepustowosc PCIe. Space complexity: VRAM ~ (warstwy na GPU / wszystkie warstwy) x rozmiar modelu.

Benchmark notes

PCIe 4.0 x16 bandwidth is ~32 GB/s, versus ~1-3 TB/s VRAM bandwidth of modern GPUs - so each offloaded layer significantly slows generation. In practice llama.cpp lets you smoothly tune n-gpu-layers: the more layers fit in VRAM, the higher the tokens/s; full CPU offload reduces speed to CPU-inference levels.

Compute bottleneck

PCIe transfer

The bottleneck is the PCIe bus bandwidth between CPU/disk and GPU; it, not compute, limits speed under offload.

Depends on
Liczba offloadowanych warstwGeneracja PCIe / NVLink

Execution paradigm

Primary mode
Dense

All layers execute; the difference is the memory location of the weights.

Activation pattern
All paths active
Routing mechanism

Offloading does not change the compute graph or introduce routing; it only controls where weights are stored and transferred.

Parallelism

Parallelism level
Conditionally parallel

Transfer and compute can overlap (prefetching), but inter-layer dependency forces sequential activation flow; cross-device parallelism is conditional on the bus.

Scope
InferenceAcross devicesAcross layers

Hardware requirements

Primary

The GPU runs the layer compute, but the benefit of offload depends on a fast bus (PCIe 4.0/5.0, NVLink) and large, fast CPU memory.

Good fit

Offloaded layers can be computed directly on the CPU (e.g. hybrid mode in llama.cpp) when transferring to the GPU is not worthwhile for a given layer.