AI Accelerator · serves as: AI acceleration, AI Inference, Compute.
Which group Cerebras Wafer-Scale Engine belongs to and how it is built
This subcategory groups integrated circuits designed to accelerate AI computation in data centers: NVIDIA H100/B200, Google TPU, AWS Trainium/Inferentia, Meta MTIA, Groq LPU, Cerebras WSE. Common properties: high-bandwidth HBM memory, low-precision modes (FP16/BF16/FP8/MX8/MX4/INT8/INT4), scalability to rack-scale (hundreds of chips in a single scale-up domain), liquid cooling, integration with ML frameworks (PyTorch, JAX, Triton, vLLM). This contrasts with hardwareSubcategory.ai-soc-edge-ai-soc where the criteria are the opposite (low power, no HBM, on-device inference).
An AI Accelerator is a specialized hardware component designed for efficient execution of artificial intelligence computations, particularly neural network inference, computer vision processing, and sensor data analysis. In robotics, AI accelerators are used to run perception models, object recognition, image segmentation, planning, and other tasks that require high computational throughput under constrained power budgets. They may take the form of dedicated NPU, TPU, VPU, or GPU chips, or specialized embedded modules.
A wafer-scale AI processor is a design class where, instead of dicing a wafer into many separate dies and connecting them in a system, the entire silicon wafer remains a single monolithic processor with an area on the order of tens of thousands of mm². It is characterised by: a huge number of cores (hundreds of thousands) and transistors (trillions); very large on-wafer SRAM (tens of GB) with petabytes-per-second bandwidth; a wafer-scale fabric connecting the cores at petabytes/s; the elimination of chip-to-chip communication bottlenecks; and special packaging, (liquid) cooling and power delivery. The class describes physical construction, not the functional role (which is given by the 'AI Accelerator' type). The flagship example is the Cerebras Wafer-Scale Engine (WSE-1/2/3).
About 4 trillion.
AI-optimized cores.
Wafer-scale interconnect.
Peak (FP16).
Powers the Cerebras CS-3 system.
Peak (FP16).
Powers the Cerebras CS-3 system.
Other hardware parts related to Cerebras Wafer-Scale Engine
The Cerebras Wafer-Scale Engine (WSE) is an AI processor developed by Cerebras Systems in which an entire silicon wafer forms a single giant integrated circuit — unlike the traditional approach where a wafer is diced into many smaller chips. This lets the WSE eliminate chip-to-chip communication bottlenecks and offer huge memory and bandwidth directly on the wafer.
The latest generation, WSE-3 (2024), is built on the TSMC 5 nm process, measures 46,225 mm² and contains 4 trillion transistors and 900,000 AI-optimized cores. It offers 44 GB of on-wafer SRAM, 21 PB/s of memory bandwidth and 27 PB/s of wafer-scale fabric bandwidth, reaching 125 petaFLOPS of AI compute. According to Cerebras that is about 19x more transistors and 28x more compute than the NVIDIA B200.
The WSE powers the Cerebras CS systems (currently CS-3), used for training and fast inference of large language models. Earlier generations are WSE-1 (2019: 1.2T transistors, 400,000 cores, 18 GB SRAM) and WSE-2 (2021: 2.6T transistors, 850,000 cores, 40 GB SRAM, TSMC 7 nm).
