Our take
Newless than six months old and changing fast- Activity
- 80
- Community
- 74
- Issues
- 77
- Pull requests
- 85
- Merges more pull requests than 97% of the projects we track
- Closes issues faster than 83% of the projects we track
- More contributors than 77% of the projects we track
colibri is an ambitious, actively developed inference project for people willing to trade speed and setup simplicity for access to large local MoE models.
Good fit if
- You want to run supported open MoE models on hardware you already own, including CPU-only systems.
- You need a local inference server or dashboard and can work with its Python-based setup tools.
- You are interested in measuring or tuning expert placement, storage, and hardware backends.
Look elsewhere if
- You need consistently fast generation on modest hardware, since disk-streamed inference can be slow.
- You need a mature, slowly changing dependency: the project is new and its development and releases are moving quickly.
- You need every GPU path to have independently measured performance, since the README says discrete-GPU Vulkan performance has not yet been measured by the project.
All health signals
| Last commit | 2026-10-06 (today) |
|---|---|
| Commits, last 90 days | 100+ |
| Releases, last 12 months | 20 (latest v2.0.0, 2026-10-06) |
| Contributors | 100+ (top contributor: 59% of commits) |
| Issues closed, last 90 days | 5+ (typically closed in 0 days) |
| Project age | 3 months |
Checked on 2026-10-06 with the GitHub API.
Overview
colibri runs large open mixture-of-experts models by keeping the dense parts in RAM and streaming routed experts from storage as needed. This targets a practical bottleneck: model files can be much larger than a computer's fast memory, even though only a fraction of the experts are used for each token.
It is a good option for developers and researchers who want to experiment with local inference across a wide range of hardware, from CPU-only machines to systems with GPUs. The tradeoff is that disk streaming can make generation slow, especially on large models or slower drives.
Key Features
- Written in C, with no GPU required for CPU inference.
- Streams mixture-of-experts weights from disk and caches frequently used experts in RAM or GPU memory.
- Supports multiple model families, including GLM, DeepSeek, Qwen, MiMo, Kimi, and others listed in the README.
- Offers Vulkan, CUDA, and Apple Silicon Metal paths for supported engines and hardware.
- Includes a terminal chat client, browser dashboard, and local server with OpenAI- and Anthropic-compatible APIs.
- Provides a System One API for structured decisions with probabilities over specified choices.
- Includes setup scripts that inspect hardware, recommend models, and resume interrupted downloads.
Use Cases
- Run an open MoE model locally when its full weights do not fit in RAM, and slower disk-bound generation is acceptable.
- Prototype local AI features in an application using the OpenAI-compatible API, or connect tools that support the Anthropic API.
- Compare CPU, GPU, storage, and cache configurations while researching inference performance.
- Use structured model scoring for bounded choices, such as routing or triage, when the allowed answers and their probabilities matter more than generated prose.
What you need
Detected in the repository
- Python >=3.10 (from pyproject.toml)
- Automated checks on GitHub Actions
License in plain words
Apache-2.0permissive
- Commercial use: yes
- Modify and redistribute: yes
- You must keep: the license, the NOTICE file and a note of your changes
- Share your changes: no
- Includes an explicit patent grant from the contributors:
A summary, not legal advice: the LICENSE file is what applies.
Getting Started
On Linux, the README's quick setup is:
sudo apt install git python3 build-essential
git clone https://github.com/JustVugg/colibri
cd colibri
./start-here.sh
The setup recommends a model based on the machine and downloads it. See the README for Windows, macOS, manual installation, and model requirements.
Alternatives
Considerations
- The project is new and changes quickly. Its recent activity and release cadence are strong, but teams should expect behavior and supported-model details to evolve.
- Disk capacity and drive speed matter: even the smallest listed model needs substantial storage, and streamed experts can constrain throughput. The README recommends at least 8 GB RAM and more for a better experience.
- Python 3.10 or newer is required by the project tooling. GPU acceleration is optional, but choosing a GPU backend can involve additional drivers or toolkits.
- The repository is Apache-2.0 licensed, but model weights retain their own licenses. The README specifically notes that Qwen-Image-2.1 is non-commercial, so check each model's terms before use.
- The repository's health data does not report a test suite, although the README describes CI checks and a tests directory. Verify the project's checks against your own adoption requirements.
Found this useful?
Share it with someone who would like colibri.