Vulkan
How it works
Vulkan exposes the GPU through explicit objects: instances, physical and logical devices, queues, command buffers, pipelines, and descriptor sets binding resources. The programmer manages memory and synchronization directly (barriers, semaphores, fences), giving low overhead and predictable performance at the cost of more complexity than older APIs. General-purpose compute uses compute shaders, usually written in GLSL/HLSL and compiled to the SPIR-V intermediate representation, launched in workgroups massively in parallel. In AI this is used to implement vendor-independent kernels (matrix multiply, quantized kernels); for example, the Vulkan backend in llama.cpp runs LLM inference on AMD, Intel and NVIDIA cards without CUDA. On Apple platforms Vulkan runs through the MoltenVK layer that translates to Metal.
Problem solved
GPU AI acceleration was long tied to closed, vendor-specific stacks (chiefly CUDA on NVIDIA), which hindered portability to AMD, Intel or mobile hardware. Vulkan solves this as an open, vendor-independent standard: one set of compute shaders can run on GPUs from different vendors and systems (Windows, Linux, Android, and via MoltenVK also macOS/iOS), while providing low driver overhead and explicit, multi-threaded control over GPU work scheduling.
Key mechanisms
Strengths & limitations
Components
General-purpose compute programs launched in workgroups on the GPU, used to implement vendor-independent AI kernels.
A portable binary intermediate format that shaders are compiled to (from GLSL/HLSL), enabling them to run across different drivers.
Command buffers, queues, barriers, semaphores and fences giving the programmer full, multi-threaded control over GPU work and memory at low overhead.
Implementation
Explicit memory management and synchronization make Vulkan harder and more error-prone than higher-level APIs.
Performance and implementation completeness depend on the vendor's driver; the same shaders can run differently on GPUs of different brands.
Evolution
Vulkan 1.0 debuts as the successor to OpenGL: an open, low-overhead, cross-platform graphics and compute standard.
Successive versions expand compute shaders, subgroups and interoperability, strengthening Vulkan as a GPGPU platform.
llama.cpp gains a Vulkan backend, enabling LLM inference on AMD, Intel and NVIDIA GPUs without CUDA, popularizing Vulkan in local AI.
Hyperparameters (configurable axes)
Number of threads in a compute shader's workgroup, affecting GPU utilization.
The subgroup (warp/wavefront) size used in subgroup operations for performance.
The GPU and driver for which SPIR-V shaders are compiled and optimized.
Computational complexity
Time complexity: Zależna od jądra (np. O(N^2 d) dla uwagi) — Vulkan to warstwa wykonania. Space complexity: Ograniczona pamięcią GPU (VRAM) urządzenia docelowego.
The Vulkan backend in llama.cpp enables LLM inference on a wide range of GPUs (AMD, Intel Arc, NVIDIA) without CUDA/ROCm, which is crucial on hardware where vendor stacks are unavailable or immature. Performance can be lower than native CUDA on NVIDIA cards, but portability makes Vulkan an attractive universal backend, especially for Intel and AMD GPUs in local inference.
Compute bottleneck
Vulkan performance depends on the vendor's driver implementation and compute-shader optimization; the standard itself has low overhead.
Execution paradigm
The activation pattern is set by the model/kernel being run, not by the Vulkan API itself.
Vulkan is a compute-shader execution API; it imposes no routing or conditional activation (those come from the model).
Parallelism
Vulkan compute shaders execute massively in parallel in workgroups; the degree of parallelism depends on the kernel and GPU.
Hardware requirements
Vulkan runs on GPUs from all major vendors (AMD, NVIDIA, Intel, ARM Mali), exposing their compute units through compute shaders.
As an open Khronos standard, Vulkan is portable across systems and hardware; on Apple it runs via the MoltenVK layer translating to Metal.