Metal
How it works
Metal exposes the GPU through objects: a device (MTLDevice), command queues/buffers, compute pipeline states, and resources (buffers and textures). GPU programs are written in the Metal Shading Language (C++-based), compiled into kernels launched massively in parallel across GPU cores. For AI the key pieces are Metal Performance Shaders (MPS) and MPSGraph — optimized kernel libraries (matrix multiply, convolution, softmax, attention) that framework backends build on. Thanks to Apple silicon's unified memory, the same memory pages are accessible to CPU and GPU, so tensors need no explicit transfer; this especially benefits LLMs, where a large KV cache and weights can reside in shared high-bandwidth memory.
Problem solved
Running and training AI models on Apple hardware requires direct, efficient GPU access with minimal driver overhead. Earlier APIs (OpenGL/OpenCL) had high overhead and were deprecated by Apple. Metal solves this as a native, low-level layer that lets AI frameworks run tensor operations on the Apple silicon GPU, leveraging unified memory (shared CPU and GPU memory) without costly cross-device data copies.
Key mechanisms
Strengths & limitations
Components
A C++-based language for writing shaders and compute kernels compiled to run on the Apple GPU.
Optimized kernel libraries for linear algebra, convolutions and compute graphs that AI backends rely on.
The mechanism for queuing and submitting work to the GPU; encodes compute and render commands for asynchronous execution.
Implementation
Not all AI framework operators have native Metal/MPS kernels; missing ops can fall back to CPU, hurting performance.
Metal-based code is not portable to other platforms (Windows, Linux, NVIDIA/AMD GPUs), complicating cross-platform deployment.
Evolution
Metal debuts on iOS as a low-overhead graphics and compute API replacing OpenGL ES on Apple hardware.
Apple adds MPS with optimized kernels, laying the groundwork for machine-learning acceleration.
The Mac's transition to Apple silicon (M1) with unified memory makes Metal the main AI acceleration layer on Macs, used by PyTorch MPS, MLX and llama.cpp.
Hyperparameters (configurable axes)
Number of threads in a compute kernel's threadgroup, affecting GPU utilization.
Kernel data type: FP32, FP16 (half) or INT.
Computational complexity
Time complexity: Zależna od jądra (np. O(N^2 d) dla uwagi) — Metal to warstwa wykonania. Space complexity: Ograniczona pojemnością unified memory (wspólna CPU/GPU).
Apple silicon unified memory (e.g. M2 Ultra up to 192 GB, M-series with hundreds of GB/s bandwidth) allows large LLMs to run locally without a dedicated GPU, because weights and the KV cache fit in shared memory. llama.cpp and MLX use Metal for LLM inference on Macs with tokens/s competitive against mid-range discrete GPUs, at low power draw.
Compute bottleneck
Metal's AI performance depends on GPU core count and unified-memory bandwidth; the interface itself has low overhead.
Execution paradigm
The activation pattern is set by the model being run, not by the Metal API itself.
Metal is a kernel-execution API; it imposes no routing or conditional activation (those come from the model being run).
Parallelism
Metal launches kernels massively in parallel across GPU cores; the degree of parallelism depends on the kernel and hardware.
Hardware requirements
Metal is the native GPU API for Apple silicon (M-series) and exploits unified memory and the GPU's hardware compute units for AI acceleration.
Metal is limited to Apple platforms (macOS, iOS, iPadOS); it does not run on hardware or systems outside the Apple ecosystem.