llama-cpp-python: Run llama.cpp Models from Python

Summary
Python bindings for llama.cpp let developers run GGUF language models locally through a high-level API, low-level C bindings, or an OpenAI-compatible server. Useful when you want local inference and control over hardware backends.
At a glance
- Language
- Python
- License
- MIT
- Stars
- 10.6k
- Forks
- 1.5k
- Added to OSRepos
- November 11, 2025
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
llama-cpp-python connects Python applications to llama.cpp, enabling local inference with compatible models. It offers both a managed Python interface and direct access to llama.cpp’s C API, so projects can choose convenience or lower-level control.
It is aimed at developers building local language-model features or serving models behind an OpenAI-compatible API. Unlike a hosted model service, it runs the model on the user’s machine, with performance and setup depending on the selected model, hardware, and build configuration.
Key Features
- Generate text and chat completions with the
LlamaPython API. - Access the underlying llama.cpp C API through Python
ctypesbindings. - Run an OpenAI-compatible web server for existing API clients.
- Configure CPU and hardware acceleration backends, including CUDA, Metal, HIP, and Vulkan.
- Support structured JSON output, JSON Schema, and tool or function calling.
- Work with supported multimodal models, including image input.
- Create embeddings and use speculative decoding.
- Load compatible GGUF models from local files or the Hugging Face Hub.
Use Cases
- Python developers can add local text or chat generation to an application without sending prompts to a hosted inference API.
- Teams can expose a local GGUF model through an OpenAI-compatible endpoint for clients that already use that API style.
- Researchers and engineers can experiment with model settings, embeddings, or low-level llama.cpp functions from Python.
- Developers can build image-and-text workflows with supported multimodal models and chat handlers.
Project Facts
- Language: Python
- License: MIT
- Stars: 10.6k
- Forks: 1.5k
- Topics: none listed
- Archived: no
Getting Started
Install the package with pip:
pip install llama-cpp-python
The default installation builds llama.cpp from source. See the README for model examples, hardware-specific build options, and server setup.
Alternatives
- textgen: TextGen offers a user-facing web interface and model management, while llama-cpp-python provides Python bindings and an OpenAI-compatible server for llama.cpp.
- LocalAI: LocalAI is a multi-backend server for language, vision, audio, and image models, while llama-cpp-python focuses on llama.cpp inference from Python and its compatible server.
Considerations
- Installation requires Python 3.8 or later and a C compiler. Building from source may require configuring CMake and the desired acceleration backend.
- Inference is local, so available memory and hardware constrain which models and context sizes are practical. GPU acceleration requires a compatible build and device.
- The Python API depends on llama.cpp, so changes to its C API may require corresponding binding updates.
- The repository is actively developed, and the README describes installation and supported options that can vary by platform and backend.
Found this useful?
Share it with someone who would like llama-cpp-python.
Comparisons
Source repository
Open the original repository on GitHub.
12 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

agentevals: Evaluate AI Agents from OpenTelemetry Traces
October 4, 2026
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.