GGML
Pure C/C++ tensor library for machine learning enabling efficient inference of large models on commodity hardware (CPU/GPU); the foundation of llama.cpp and whisper.cpp.

Description
GGML is a machine-learning tensor library written in plain C/C++ with no third-party dependencies. It was created by Georgi Gerganov (work began in September 2022) and is developed within the ggml-org organization. It is designed to run large models with high performance on commodity hardware.
Key features include 2- to 8-bit integer quantization plus MXFP4 and NVFP4 microscaling formats, zero memory allocations during runtime, SIMD-optimized kernels (x86, ARM, RISC-V), and broad backend support: CPU, GPU (CUDA, Metal, Vulkan), NPU, and the browser (WebAssembly). The library is cross-platform (x86, ARM, RISC-V, LoongArch, PowerPC, s390x, WASM).
GGML is the computational foundation of the llama.cpp and whisper.cpp projects. Models are stored in the binary GGUF format (GGML Universal File, introduced in August 2023), which superseded the earlier GGML format. The library is released under the MIT license.