Ultrafast Inference
How it works
The layer intercepts incoming requests and merges them on the fly (continuous batching), adding new sequences to the running batch instead of waiting for prior ones to finish. Attention state is kept in a KV-cache, often managed in pages (PagedAttention) to limit memory fragmentation. Speculative decoding uses a small draft model to propose several tokens at once, which the large model verifies in parallel. Quantization (e.g. to 8 or 4 bits) reduces memory and bandwidth usage. Optimized attention kernels minimize data movement between HBM and registers. Everything runs on tensor-core GPUs or dedicated accelerators, frequently with tensor and pipeline parallelism across devices.
Problem solved
Naive LLM inference is slow and expensive: autoregressive generation produces one token at a time, the KV memory grows linearly with context length, and static request batching wastes GPU capacity under variable sequence lengths. An ultrafast inference layer addresses high latency and low hardware utilization when serving models in production.
Components
Dynamically adding and removing sequences from the running batch during generation, without waiting for all requests to finish.
Official
Storing and reusing key-value tensors from prior decoding steps; often managed in pages (PagedAttention) to limit memory fragmentation.
Official
A small draft model proposes several tokens at once, which the large model verifies in a single pass, accepting correct prefixes.
Official
Representing weights and/or activations in lower precision (e.g. INT8, FP8, INT4) to reduce memory and bandwidth usage.
Official
IO-aware attention kernels (e.g. FlashAttention) that minimize data movement between HBM and registers.
Official
Use of tensor-core GPUs or dedicated inference accelerators, often with tensor and pipeline parallelism across devices.
Official
Implementation
Increasing batch size raises throughput but can increase single-request latency.
Long contexts and many concurrent sequences can exhaust GPU memory.
Aggressive quantization (e.g. INT4) can reduce model output quality.
Poor alignment between draft and target model limits or negates the speedup.
Evolution
Optimized exact-attention kernels reducing memory traffic, a foundation for fast GPU inference.
Speeding up autoregressive generation via parallel verification of tokens proposed by a small model.
Paged KV-cache management and continuous batching in an open serving engine, popularizing high-throughput LLM serving.
Hyperparameters (configurable axes)
Maximum number of sequences served concurrently; a throughput-versus-latency trade-off.
Numeric precision of weights/activations (e.g. FP16, FP8, INT8, INT4).
Maximum context length; directly determines KV-cache memory usage.
Number of tokens proposed at once by the draft model; affects the gain from parallel verification.
Parallelism
Parallelism across requests (batching) and across devices (tensor/pipeline parallelism); autoregressive decoding remains sequential per sequence, partially mitigated by speculative decoding.
Hardware requirements
Tensor cores and high HBM bandwidth are crucial for matrix multiplies and low-latency KV-cache handling.
TPUs handle dense inference matrix operations well given suitable serving software.
Feasible for small models or engines like llama.cpp, but throughput and latency are worse than on accelerators.