llama.cpp
llama.cpp is an open source C/C++ inference engine for running large language models on local computers and other devices. It supports CPU and GPU execution, model quantization, and a range of hardware, helping reduce memory use and reliance on hosted services. This makes it useful for private, offline, or resource-constrained applications, as well as for experimenting with model behavior and deployment.
Tools in this area include inference runtimes, language bindings, server interfaces, and integrations for applications and agents. When choosing one, consider supported model formats and backends, hardware requirements, performance, license, maintenance activity, and compatibility with existing software. These tools can suit developers building local AI features, researchers testing models, and users who want more control over where inference runs.
2 repositories · updated September 3, 2026

agent.cpp: Building Local LLM Agents with C++ and llama.cpp
agent.cpp provides essential building blocks for developing local AI agents using C++. It leverages llama.cpp to enable efficient execution of small language models directly on your hardware. This library offers a modular approach with features like agent loops, callbacks, tools, and grammar-constrained output, making it ideal for creating custom, privacy-focused agent solutions.

llama-cpp-python: Python Bindings for llama.cpp
llama-cpp-python provides robust Python bindings for the popular llama.cpp library, enabling efficient local inference with large language models. It offers a high-level API compatible with OpenAI's API, facilitating easy integration into existing applications. The project also includes a powerful web server for local deployment and supports various hardware acceleration backends.