A neocloud builds and operates large accelerator clusters (e.g. NVIDIA H100/H200/Blackwell) connected by fast, low-latency networking with high-throughput storage. Compute is offered as bare-metal, VMs, Kubernetes or a simple API/job queue, usually billed per GPU-hour or via capacity reservations. The stack is optimized for AI workloads (drivers, libraries, distributed-training orchestration, inference serving) rather than a full enterprise service catalog.
General-purpose hyperscalers can be expensive, have long queues for the latest GPUs and a broad but not AI-optimized stack. Neoclouds address the availability and cost of dense compute for training/inference.
Large GPU pools (e.g. H100/H200/Blackwell) as the core resource.
Official
InfiniBand/NVLink linking nodes for distributed training.
Official
Bare-metal / VM / Kubernetes / API billed per GPU-hour.
Official
Heavy dependence on GPU supply and pricing (mostly NVIDIA).
GPU purchases financed by debt create risk if demand falls.
Narrower services and certifications than hyperscalers.
Early GPU-focused providers for ML (e.g. Lambda, CoreWeave).
Surging GPU demand pushes neoclouds into the mainstream.
Large contracts and data-center buildouts (incl. for projects like Stargate).
Neoclouds are built around dense GPU clusters.