mlx-lm
A Python package for running and fine-tuning LLMs on Apple silicon, built on the MLX framework — with quantization, Hugging Face model conversion, and an OpenAI-compatible server.

Description
mlx-lm is a Python package for running (inference) and fine-tuning large language models (LLMs) on computers with Apple silicon. It is developed by Apple's machine learning research team (the ml-explore organization on GitHub) and released under the MIT license.
The package is built on the MLX framework — it is a separate project that uses MLX as its compute layer. It provides integration with the Hugging Face Hub (downloading and publishing models), model quantization and conversion, full and LoRA fine-tuning (including for quantized models), and distributed inference and training via mx.distributed.
mlx-lm exposes a command-line interface (including mlx_lm.generate, mlx_lm.chat, mlx_lm.convert, mlx_lm.lora) and a Python API. It also includes an HTTP server with an endpoint compatible with the OpenAI chat API, allowing locally run models to be used with tools built for OpenAI.