llama.cpp
C/C++ inference engine for LLMs on CPUs and GPUs, using the GGUF format and quantization.
Overview
C/C++ inference engine for LLMs on CPUs and GPUs, using the GGUF format and quantization.
- Category
- LLM runtimes · AI tools
- Platforms
- Linux, macOS, Windows, API only
- Pricing
- Free
- License
- MIT
- Source code
- ggml-org/llama.cpp
- Repository stars
- 128,722
- API
- No public API description
Release history
Updated 3 days ago10 releases in the last 12 months.
-
b11030
-
b11033
-
b11034
-
b11035
-
b11036
-
b11037
What users say
No reviews yet
Ratings appear once verified reviews arrive. Until then the rank uses public signals only: release freshness and repository adoption.
Alternatives to llama.cpp
See all alternativesRuntime to download and run large language models locally, with a CLI and a REST API.
High-throughput LLM serving engine with paged attention and an OpenAI-compatible server.
Self-hosted, OpenAI-compatible API that runs text, image and audio models on local hardware.