Robots Atlas>ROBOTS ATLAS
Infrastructure

Neocloud

2023ActiveUpdated: 19 August 2026Published
Key innovation
A specialized cloud provider built from the ground up around AI accelerators (GPUs/โ€ฆ), offering dense compute for training and inference more cheaply and readily than general-purpose hyperscalers.
Category
Infrastructure
Abstraction level
Paradigm
Operation level
DeploymentServing
Use cases
Training large models (LLMs, multimodal models)Large-scale inferenceFine-tuning and research experimentsAccess to the latest GPUs without long waitsCapacity reservations for AI labs

How it works

A neocloud builds and operates large accelerator clusters (e.g. NVIDIA H100/H200/Blackwell) connected by fast, low-latency networking with high-throughput storage. Compute is offered as bare-metal, VMs, Kubernetes or a simple API/job queue, usually billed per GPU-hour or via capacity reservations. The stack is optimized for AI workloads (drivers, libraries, distributed-training orchestration, inference serving) rather than a full enterprise service catalog.

Problem solved

General-purpose hyperscalers can be expensive, have long queues for the latest GPUs and a broad but not AI-optimized stack. Neoclouds address the availability and cost of dense compute for training/inference.

Key mechanisms

Dense GPU clusters with fast networking (InfiniBand/NVLink)
Per-GPU-hour billing / capacity reservations
Bare-metal or a simple API instead of a full cloud stack
An AI-optimized stack (distributed training, serving)
A focus on availability of the latest accelerators

Strengths & limitations

Strengths
โœ“Lower unit cost of GPU compute
โœ“Faster access to the latest accelerators
โœ“AI-optimized performance
โœ“Flexible billing models
Limitations
โœ—Narrower service scope than hyperscalers (fewer managed services)
โœ—Dependence on GPU supply and debt financing
โœ—Less mature enterprise/compliance features
โœ—Customer-concentration risk and GPU-market cyclicality

Components

Accelerator clusterCompute resource

Large GPU pools (e.g. H100/H200/Blackwell) as the core resource.

Official

High-speed interconnectCommunication

InfiniBand/NVLink linking nodes for distributed training.

Official

Access layerProvisioning

Bare-metal / VM / Kubernetes / API billed per GPU-hour.

Official

Implementation

Implementation pitfalls
Single hardware-vendor dependenceMedium

Heavy dependence on GPU supply and pricing (mostly NVIDIA).

Fix:Diversify suppliers, use reservations, plan capacity.
Debt-heavy financial modelHigh

GPU purchases financed by debt create risk if demand falls.

Fix:Long-term contracts, customer diversification.
Enterprise/compliance gapsMedium

Narrower services and certifications than hyperscalers.

Fix:Add certifications and managed services.

Evolution

2017
Specialized GPU clouds emerge

Early GPU-focused providers for ML (e.g. Lambda, CoreWeave).

2023
Generative-AI boom and GPU shortage
Inflection point

Surging GPU demand pushes neoclouds into the mainstream.

2025
Neoclouds as an AI-infrastructure pillar

Large contracts and data-center buildouts (incl. for projects like Stargate).

Hyperparameters (configurable axes)

Accelerator typeHigh
NVIDIA H100/H200/BlackwellThe most common GPUs in neoclouds.
inne (AMD, ASIC)Less often โ€” AMD Instinct, ASIC accelerators.
Access modelMedium
bare-metal / APIBare-metal, VMs, Kubernetes or a simple API.
BillingMedium
per GPU-hour / reservationPer GPU-hour or capacity reservation.

Hardware requirements

Primary

Neoclouds are built around dense GPU clusters.