Accelerators combine many specialized compute units (e.g. Tensor Cores, systolic arrays) operating in parallel, support reduced precisions (FP16/BF16/FP8/FP4/INT8) and fast high-bandwidth memory (HBM) to maximize operations per watt and per second.
General-purpose CPUs are inefficient for the massively parallel tensor operations of neural networks, limiting speed and raising the energy cost of training and inference.