Robots Atlas>ROBOTS ATLAS
GoogleDeployed

Google TPU Trillium (v6e)

Trillium is Google’s sixth-generation AI accelerator (TPU v6e), a dedicated ASIC for accelerating neural-network computation in training and inference. A single chip delivers peak performance of 918 TFLOPs in bfloat16 and 1836 TOPS in int8, 32 GB of HBM with 1638 GB/s bandwidth, and an 800 GB/s bidirectional inter-chip interconnect (ICI) across 4 ports. Each chip contains one TensorCore with two matrix-multiply units (MXUs).

Chips are combined into pods of 256 (2D torus topology, up to a 16x16 configuration), with 8 chips per host. Google offers them as Cloud TPU in VM configurations from 1 to 8 chips. Within a Jupiter network fabric up to 100,000 Trillium chips can be connected (13 Pb/s bisectional bandwidth), with 99% scaling efficiency across 12 pods (3,072 chips) and 94% across 24 pods (6,144 chips) for GPT3-175B pretraining.

Compared with the previous v5e generation, Trillium provides 4.7x higher peak compute per chip, up to 4x faster training and up to 3x higher inference throughput, double the HBM capacity and bandwidth and double the ICI bandwidth, at 67% better energy efficiency. Trillium became generally available (GA) on December 12, 2024.

#AI accelerator#ASIC#TPU#Trillium#v6e#bfloat16#Google Cloud#machine learning#inference
Market classIndustrial
Introduction date12 December 2024

Classification

AI Accelerator · serves as: AI acceleration, AI Inference, Compute.

Which group Google TPU Trillium (v6e) belongs to and how it is built

5items
Compute
Hardware domain
Industrial
Market class

This subcategory groups integrated circuits designed to accelerate AI computation in data centers: NVIDIA H100/B200, Google TPU, AWS Trainium/Inferentia, Meta MTIA, Groq LPU, Cerebras WSE. Common properties: high-bandwidth HBM memory, low-precision modes (FP16/BF16/FP8/MX8/MX4/INT8/INT4), scalability to rack-scale (hundreds of chips in a single scale-up domain), liquid cooling, integration with ML frameworks (PyTorch, JAX, Triton, vLLM). This contrasts with hardwareSubcategory.ai-soc-edge-ai-soc where the criteria are the opposite (low power, no HBM, on-device inference).

An AI Accelerator is a specialized hardware component designed for efficient execution of artificial intelligence computations, particularly neural network inference, computer vision processing, and sensor data analysis. In robotics, AI accelerators are used to run perception models, object recognition, image segmentation, planning, and other tasks that require high computational throughput under constrained power budgets. They may take the form of dedicated NPU, TPU, VPU, or GPU chips, or specialized embedded modules.

A data-center AI accelerator card is a design class describing the construction of high-performance compute processors (GPUs/accelerators) intended for mounting in data-center servers. It is characterised by: an SXM form factor (a module soldered onto an HGX/DGX baseboard) or a dual-slot PCIe card; high-bandwidth memory (HBM2e/HBM3/HBM3e) integrated on-package; dedicated GPU-to-GPU interconnects (NVLink, Infinity Fabric) with hundreds of GB/s of bandwidth; high TDP (350–1000 W) requiring air or liquid cooling; support for virtualisation/partitioning (MIG) and low-precision compute formats (FP8/FP16/BF16/INT8). The class includes designs such as NVIDIA H100/H200/A100, AMD Instinct MI300, Google TPU, and Intel Gaudi. It describes physical construction and configuration, not the functional role (which is given by the component type "AI Accelerator").

Technical specification

Compute performance
3parameters
Peak compute (bf16)
918 TFLOPs

per chip

Peak compute (Int8)
1836 TOPS

per chip

TensorCore / MXU
1 TensorCore, 2 MXU

per chip

Updated: 29 September 2026