Robots Atlas>ROBOTS ATLAS
Inference

LPU

2020ActiveUpdated: 19 August 2026Published
Key innovation
A deterministic dataflow processor (Tensor Streaming Processor) designed by Groq exclusively for language-model inference: no HBM and no reactive hardware, with on-chip SRAM and a fully compiler-scheduled execution — yielding very low and predictable latency.
Category
Inference
Abstraction level
Paradigm
Operation level
InferenceServing
Use cases
Low-latency LLM inference (interactive chat, agents)High-token-throughput model servingApplications requiring predictable latency (real-time)Inference cloud (GroqCloud) and on-premise deployments (GroqRack)

How it works

Groq's compiler schedules every operation and data transfer in time ahead of execution (static, deterministic scheduling), so the chip needs no caches or branch prediction. Data streams through arrays of compute units (matrix multiplication, vector functions) in a fixed cadence, and weights and activations reside in fast on-chip SRAM. For large models, many LPUs are connected in a deterministically communicating network, allowing the model to scale while preserving predictable latency. The result is very fast, repeatable token generation.

Problem solved

LLM inference on GPUs is bounded by memory (HBM) bandwidth and suffers from variable latency due to reactive, non-deterministic execution. The LPU addresses this with deterministic dataflow and on-chip SRAM, delivering predictably low latency and high throughput.

Key mechanisms

Tensor Streaming Processor (TSP) — a dataflow architecture
Fully deterministic, compiler-scheduled execution
On-chip SRAM instead of HBM
No caches or branch prediction
A deterministic multi-chip network for large models

Strengths & limitations

Strengths
✓Very low and predictable latency
✓High token throughput
✓Determinism aiding optimization and SLAs
✓Eliminating the HBM bottleneck in decoding
Limitations
✗Small per-chip SRAM → large models need many chips
✗Specialized for inference (not training)
✗Dependence on the compiler and static scheduling
✗A narrower ecosystem than CUDA/GPUs

Components

Tensor Streaming Processor (TSP)Compute core

The dataflow core executing operations in a deterministic, scheduled cadence.

On-chip SRAMMemory

Large, fast memory for weights and activations instead of HBM.

Compiler (static scheduler)Execution control

Schedules every operation and data transfer ahead of time.

Deterministic chip networkScaling

Links many LPUs for large models with predictable communication.

Official

Implementation

Implementation pitfalls
Limited per-chip SRAMMedium

Large models require sharding across many LPUs.

Fix:Multi-chip networks, model partitioning, quantization.
Compiler dependenceMedium

Performance and correctness depend on the compiler's static schedule.

Fix:A mature toolchain, testing and profiling.

Evolution

Original paper · 2020 · ISCA 2020 · Dennis Abts
Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads
Dennis Abts, i in. (Groq)
2016
Groq founded

Jonathan Ross (co-creator of Google's TPU) founds Groq.

2020
TSP architecture published (ISCA)
Inflection point

Groq describes the Tensor Streaming Processor — a deterministic dataflow processor.

2024
LPU in the GroqCloud service
Inflection point

LLM inference on the LPU goes mainstream via GroqCloud.

Hyperparameters (configurable axes)

Chip countHigh
wiele LPULarge models spread across a chip network.
PrecisionMedium
FP16/INT8The compute format affects throughput and memory.

Execution paradigm

Primary mode
Dense

Execution is dense and fully deterministic (static compiler scheduling), with no reactive hardware mechanisms.

Activation pattern
All paths active

Parallelism

Parallelism level
Partially parallel

Parallelism across compute arrays and multiple chips; autoregressive decoding remains sequential across tokens.

Scope
InferenceAcross devices