Google TPU Trillium (v6e)
Trillium is Google’s sixth-generation AI accelerator (TPU v6e), a dedicated ASIC for accelerating neural-network computation in training and inference. A single chip delivers peak performance of 918 TFLOPs in bfloat16 and 1836 TOPS in int8, 32 GB of HBM with 1638 GB/s bandwidth, and an 800 GB/s bidirectional inter-chip interconnect (ICI) across 4 ports. Each chip contains one TensorCore with two matrix-multiply units (MXUs).
Chips are combined into pods of 256 (2D torus topology, up to a 16x16 configuration), with 8 chips per host. Google offers them as Cloud TPU in VM configurations from 1 to 8 chips. Within a Jupiter network fabric up to 100,000 Trillium chips can be connected (13 Pb/s bisectional bandwidth), with 99% scaling efficiency across 12 pods (3,072 chips) and 94% across 24 pods (6,144 chips) for GPT3-175B pretraining.
Compared with the previous v5e generation, Trillium provides 4.7x higher peak compute per chip, up to 4x faster training and up to 3x higher inference throughput, double the HBM capacity and bandwidth and double the ICI bandwidth, at 67% better energy efficiency. Trillium became generally available (GA) on December 12, 2024.

Classification
AI Accelerator · serves as: AI acceleration, AI Inference, Compute.
Which group Google TPU Trillium (v6e) belongs to and how it is built
This subcategory groups integrated circuits designed to accelerate AI computation in data centers: NVIDIA H100/B200, Google TPU, AWS Trainium/Inferentia, Meta MTIA, Groq LPU, Cerebras WSE. Common properties: high-bandwidth HBM memory, low-precision modes (FP16/BF16/FP8/MX8/MX4/INT8/INT4), scalability to rack-scale (hundreds of chips in a single scale-up domain), liquid cooling, integration with ML frameworks (PyTorch, JAX, Triton, vLLM). This contrasts with hardwareSubcategory.ai-soc-edge-ai-soc where the criteria are the opposite (low power, no HBM, on-device inference).
An AI Accelerator is a specialized hardware component designed for efficient execution of artificial intelligence computations, particularly neural network inference, computer vision processing, and sensor data analysis. In robotics, AI accelerators are used to run perception models, object recognition, image segmentation, planning, and other tasks that require high computational throughput under constrained power budgets. They may take the form of dedicated NPU, TPU, VPU, or GPU chips, or specialized embedded modules.
A data-center AI accelerator card is a design class describing the construction of high-performance compute processors (GPUs/accelerators) intended for mounting in data-center servers. It is characterised by: an SXM form factor (a module soldered onto an HGX/DGX baseboard) or a dual-slot PCIe card; high-bandwidth memory (HBM2e/HBM3/HBM3e) integrated on-package; dedicated GPU-to-GPU interconnects (NVLink, Infinity Fabric) with hundreds of GB/s of bandwidth; high TDP (350–1000 W) requiring air or liquid cooling; support for virtualisation/partitioning (MIG) and low-precision compute formats (FP8/FP16/BF16/INT8). The class includes designs such as NVIDIA H100/H200/A100, AMD Instinct MI300, Google TPU, and Intel Gaudi. It describes physical construction and configuration, not the functional role (which is given by the component type "AI Accelerator").
Technical specification
per chip
per chip
per chip
per chip
Relations
Other hardware parts related to Google TPU Trillium (v6e)

The TPU (Tensor Processing Unit) is Google's custom AI accelerator (ASIC) for machine learning, based on a systolic-array architecture (MXU, bfloat16). Available on Google Cloud; used e.g. to train the Gemini models.

NVIDIA's AI GPU architecture (2024) and the B200 chip: 208B transistors, two dies connected at 10 TB/s, TSMC 4NP, 192 GB HBM3e, 2nd-gen Transformer Engine (FP4) and 5th-gen NVLink; the basis of GB200 systems.

Amazon/AWS's custom AI accelerator (ASIC) for training and inference of AI models; NeuronCore cores, HBM memory, available in EC2 (Trn) instances. Trainium3: 144 GB HBM3e and 4.9 TB/s.