Robots Atlas>ROBOTS ATLAS
AI PlatformAI-native

llama.cpp

Open-source C/C++ library for LLM inference — supports the GGUF format, integer quantization and running models on CPU and GPU.

Producer:ggmlOn-Premises · EdgeReleased:Mar 10, 2023
llama.cpp
SDK / Languages
1c_cpp
Robotics-Ready
✓

Description

llama.cpp is an open-source library for large language model (LLM) inference written in plain C/C++ with minimal dependencies. It was created by Georgi Gerganov and development began in March 2023. The library is co-developed alongside the general-purpose GGML tensor library under the ggml-org organization.

llama.cpp uses the binary GGUF format to store model tensors and metadata and supports 1.5-bit to 8-bit integer quantization, allowing models to run under constrained memory. It supports many hardware backends — including CPU (AVX/AVX2/AVX512/AMX), CUDA (NVIDIA), Metal (Apple Silicon), Vulkan, SYCL, HIP (AMD) and MUSA — as well as hybrid CPU+GPU inference.

The project ships llama-server, an OpenAI-compatible API server with a built-in web UI. llama.cpp is released under the MIT License.

ApplicationsAI ApplicationsDomains and use cases this platform is best suited for — from RAG and fine-tuning to scientific research.

1

Developer EcosystemDeveloper EcosystemDeveloper resources: available SDKs, supported programming languages, and infrastructure features and model-deployment methods.

SDK Languages
C+C / C++
API Type
REST

SourcesDocumentation VaultCentralized hub of links to official sources, technical guides, repositories and release notes.

Data verified: Sep 29, 2026