1) A cloud provider joins the NVIDIA Cloud Partner program. 2) It designs and builds a data center to the NCP reference architecture (GPU clusters, high-bandwidth networking, storage, security). 3) The deployment is validated against NVIDIA's specifications, ensuring performance and interoperability. 4) The partner operates a GPU cloud service for customers (AI training/inference), using technical support, go-to-market and priority GPU allocation.
Building large GPU data centers for AI is complex, costly and error-prone. NCP provides a proven blueprint and NVIDIA support, reducing risk, shortening deployment time and ensuring interoperability and access to GPUs.
A validated blueprint for an AI data center (compute, networking, storage, security).
Standardized configurations with the latest NVIDIA GPUs (from ~128 to 16,000+ nodes).
Official
High-bandwidth networking (e.g. InfiniBand/Spectrum-X) and storage per spec.
Official
Technical support, go-to-market and priority GPU allocation for partners.
Official
An architecture built on NVIDIA GPUs and stack increases dependence on a single vendor.
Building clusters requires huge capital, power and contracted GPUs (constrained supply).
Maintaining compliance with an evolving reference architecture requires continual updates.
NVIDIA formalizes the Cloud Partner program and publishes a reference architecture for AI data centers, standardizing GPU cloud build-outs.
Growth of the 'neocloud' and partner ecosystem (e.g. Cisco with NCP-compliant networking) building 'AI factories' to NVIDIA's architecture.
Time complexity: Nie dotyczy (program/architektura referencyjna, nie algorytm). Space complexity: Nie dotyczy.
A partner cloud's scale is determined by GPU availability (NVIDIA allocation), the data center's power and scale-up/scale-out networking bandwidth - not a single algorithm.
From ~128 to over 16,000 nodes per the reference architecture.
InfiniBand or Ethernet (e.g. Spectrum-X) to interconnect GPUs.
NVIDIA accelerators used (e.g. Hopper, Blackwell, Rubin).
This describes the infrastructure (GPU clusters); NCP is a program/standard, not a model compute paradigm.
The whole point is scalable, parallel GPU clusters for training and inference.
The whole program is built on NVIDIA GPU clusters (tensor cores) for AI training and inference.