Robots Atlas>ROBOTS ATLAS
Infrastructure

Metal

2014ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Apple's low-level, low-overhead graphics and compute API that gives direct access to the GPU on Apple silicon (and formerly AMD/Intel on the Mac), with unified memory that eliminates CPU-GPU copies — the foundation of AI acceleration on Apple hardware.
Category
Infrastructure
Abstraction level
Building block
Operation level
InferenceServingDeployment
Use cases
PyTorch acceleration backend (MPS device) on Apple siliconThe MLX and mlx-lm framework for training and inference on MacsAccelerating LLM inference in llama.cpp (Metal backend)3D graphics, games and rendering on macOS / iOS / iPadOSGPGPU compute with unified memory without CPU-GPU copies

How it works

Metal exposes the GPU through objects: a device (MTLDevice), command queues/buffers, compute pipeline states, and resources (buffers and textures). GPU programs are written in the Metal Shading Language (C++-based), compiled into kernels launched massively in parallel across GPU cores. For AI the key pieces are Metal Performance Shaders (MPS) and MPSGraph — optimized kernel libraries (matrix multiply, convolution, softmax, attention) that framework backends build on. Thanks to Apple silicon's unified memory, the same memory pages are accessible to CPU and GPU, so tensors need no explicit transfer; this especially benefits LLMs, where a large KV cache and weights can reside in shared high-bandwidth memory.

Problem solved

Running and training AI models on Apple hardware requires direct, efficient GPU access with minimal driver overhead. Earlier APIs (OpenGL/OpenCL) had high overhead and were deprecated by Apple. Metal solves this as a native, low-level layer that lets AI frameworks run tensor operations on the Apple silicon GPU, leveraging unified memory (shared CPU and GPU memory) without costly cross-device data copies.

Key mechanisms

Metal Shading Language (C++) for shaders and kernels
Metal Performance Shaders (MPS) / MPSGraph - AI kernels
Command buffer and command queue (asynchronous execution)
Unified memory on Apple silicon (no CPU-GPU copies)
Compute pipeline state for GPGPU compute

Strengths & limitations

Strengths
✓Native, low-overhead GPU API for the whole Apple ecosystem
✓Unified memory on Apple silicon eliminates CPU-GPU copies
✓Optimized AI kernels via MPS and MPSGraph
✓Supports PyTorch MPS, MLX/mlx-lm, llama.cpp backends
✓Explicit, multi-threaded control over GPU work
Limitations
✗Limited exclusively to Apple platforms (no portability)
✗Incomplete operator coverage in some backends (CPU fallback)
✗Ecosystem lock-in complicates cross-platform deployment
✗Smaller AI tooling ecosystem than CUDA
✗Performance depends on the Apple silicon GPU generation

Components

Metal Shading Language (MSL)Defining GPU programs

A C++-based language for writing shaders and compute kernels compiled to run on the Apple GPU.

Metal Performance Shaders (MPS) / MPSGraphHigh-performance AI/ML primitives

Optimized kernel libraries for linear algebra, convolutions and compute graphs that AI backends rely on.

Command buffer and command queueScheduling GPU work

The mechanism for queuing and submitting work to the GPU; encodes compute and render commands for asynchronous execution.

Implementation

Implementation pitfalls
Incomplete operation coverage in backendsMedium

Not all AI framework operators have native Metal/MPS kernels; missing ops can fall back to CPU, hurting performance.

Fix:Check operator support in the backend (e.g. PyTorch MPS) and enable fallback or update to newer versions.
Lock-in to the Apple ecosystemLow

Metal-based code is not portable to other platforms (Windows, Linux, NVIDIA/AMD GPUs), complicating cross-platform deployment.

Fix:Use cross-platform abstractions (e.g. framework backends) or intermediate layers when portability is needed.

Evolution

Original paper · 2014 · Apple (WWDC 2014) · Apple Inc.
Metal — Apple Developer Documentation
Apple Inc.
2014
Apple unveils Metal at WWDC
Inflection point

Metal debuts on iOS as a low-overhead graphics and compute API replacing OpenGL ES on Apple hardware.

2015
Metal Performance Shaders (MPS)

Apple adds MPS with optimized kernels, laying the groundwork for machine-learning acceleration.

2020
Apple silicon and unified memory
Inflection point

The Mac's transition to Apple silicon (M1) with unified memory makes Metal the main AI acceleration layer on Macs, used by PyTorch MPS, MLX and llama.cpp.

Hyperparameters (configurable axes)

Threadgroup sizeMedium

Number of threads in a compute kernel's threadgroup, affecting GPU utilization.

256A typical group size.
Compute precisionHigh

Kernel data type: FP32, FP16 (half) or INT.

half (FP16)Faster inference, less memory.
FP32Higher accuracy.

Computational complexity

Computational characteristics
→Low-overhead GPU API with explicit resource management
→High-bandwidth unified memory (Apple silicon)
→AI kernels via MPS/MPSGraph
→Support for FP32/FP16 (half) and INT precision
→Backend for PyTorch MPS, MLX, llama.cpp

Time complexity: Zależna od jądra (np. O(N^2 d) dla uwagi) — Metal to warstwa wykonania. Space complexity: Ograniczona pojemnością unified memory (wspólna CPU/GPU).

Benchmark notes

Apple silicon unified memory (e.g. M2 Ultra up to 192 GB, M-series with hundreds of GB/s bandwidth) allows large LLMs to run locally without a dedicated GPU, because weights and the KV cache fit in shared memory. llama.cpp and MLX use Metal for LLM inference on Macs with tokens/s competitive against mid-range discrete GPUs, at low power draw.

Compute bottleneck

Apple silicon GPU and unified-memory bandwidth

Metal's AI performance depends on GPU core count and unified-memory bandwidth; the interface itself has low overhead.

Depends on
Generacja i liczba rdzeni GPUPrzepustowość unified memory

Execution paradigm

Primary mode
Dense

The activation pattern is set by the model being run, not by the Metal API itself.

Activation pattern
All paths active
Routing mechanism

Metal is a kernel-execution API; it imposes no routing or conditional activation (those come from the model being run).

Parallelism

Parallelism level
Fully parallel

Metal launches kernels massively in parallel across GPU cores; the degree of parallelism depends on the kernel and hardware.

Scope
InferenceTraining

Hardware requirements

Primary

Metal is the native GPU API for Apple silicon (M-series) and exploits unified memory and the GPU's hardware compute units for AI acceleration.

Limited

Metal is limited to Apple platforms (macOS, iOS, iPadOS); it does not run on hardware or systems outside the Apple ecosystem.