JustVugg

colibri: Run Large MoE Models on Local Hardware

colibri is a C inference engine that streams mixture-of-experts model weights from disk, letting large open models run on machines with limited RAM and no required GPU. It also provides a local dashboard and OpenAI- and Anthropic-compatible APIs.

NewCApache-2.0AILLMCLI
Stars
40k
Forks
4.4k
Last commit
today
Contributors
100+
Releases, 12 months
20

Our take

Newless than six months old and changing fast

Health score

79/100

How it is scored
Activity
80
Community
74
Issues
77
Pull requests
85
  • Merges more pull requests than 97% of the projects we track
  • Closes issues faster than 83% of the projects we track
  • More contributors than 77% of the projects we track

colibri is an ambitious, actively developed inference project for people willing to trade speed and setup simplicity for access to large local MoE models.

Good fit if

  • You want to run supported open MoE models on hardware you already own, including CPU-only systems.
  • You need a local inference server or dashboard and can work with its Python-based setup tools.
  • You are interested in measuring or tuning expert placement, storage, and hardware backends.

Look elsewhere if

  • You need consistently fast generation on modest hardware, since disk-streamed inference can be slow.
  • You need a mature, slowly changing dependency: the project is new and its development and releases are moving quickly.
  • You need every GPU path to have independently measured performance, since the README says discrete-GPU Vulkan performance has not yet been measured by the project.
All health signals
Last commit2026-10-06 (today)
Commits, last 90 days100+
Releases, last 12 months20 (latest v2.0.0, 2026-10-06)
Contributors100+ (top contributor: 59% of commits)
Issues closed, last 90 days5+ (typically closed in 0 days)
Project age3 months

Checked on 2026-10-06 with the GitHub API.

Overview

colibri runs large open mixture-of-experts models by keeping the dense parts in RAM and streaming routed experts from storage as needed. This targets a practical bottleneck: model files can be much larger than a computer's fast memory, even though only a fraction of the experts are used for each token.

It is a good option for developers and researchers who want to experiment with local inference across a wide range of hardware, from CPU-only machines to systems with GPUs. The tradeoff is that disk streaming can make generation slow, especially on large models or slower drives.

Key Features

  • Written in C, with no GPU required for CPU inference.
  • Streams mixture-of-experts weights from disk and caches frequently used experts in RAM or GPU memory.
  • Supports multiple model families, including GLM, DeepSeek, Qwen, MiMo, Kimi, and others listed in the README.
  • Offers Vulkan, CUDA, and Apple Silicon Metal paths for supported engines and hardware.
  • Includes a terminal chat client, browser dashboard, and local server with OpenAI- and Anthropic-compatible APIs.
  • Provides a System One API for structured decisions with probabilities over specified choices.
  • Includes setup scripts that inspect hardware, recommend models, and resume interrupted downloads.

Use Cases

  • Run an open MoE model locally when its full weights do not fit in RAM, and slower disk-bound generation is acceptable.
  • Prototype local AI features in an application using the OpenAI-compatible API, or connect tools that support the Anthropic API.
  • Compare CPU, GPU, storage, and cache configurations while researching inference performance.
  • Use structured model scoring for bounded choices, such as routing or triage, when the allowed answers and their probabilities matter more than generated prose.

What you need

Detected in the repository

  • Python >=3.10 (from pyproject.toml)
  • Automated checks on GitHub Actions

License in plain words

Apache-2.0permissive

  • Commercial use: yes
  • Modify and redistribute: yes
  • You must keep: the license, the NOTICE file and a note of your changes
  • Share your changes: no
  • Includes an explicit patent grant from the contributors:

A summary, not legal advice: the LICENSE file is what applies.

Getting Started

On Linux, the README's quick setup is:

sudo apt install git python3 build-essential
git clone https://github.com/JustVugg/colibri
cd colibri
./start-here.sh

The setup recommends a model based on the machine and downloads it. See the README for Windows, macOS, manual installation, and model requirements.

Alternatives

  • ds4: ds4 targets selected models on Metal, CUDA, or ROCm, while colibri streams MoE weights from disk for low-RAM systems.
  • cactus: Cactus focuses on multimodal inference on mobile and edge devices, while colibri streams MoE model weights from disk on RAM-limited machines.
ProjectLanguageLicenseStarsStatus
colibriCApache-2.040kNew
ds4CMIT23.2kNew
cactusC++–6.1kActive

Considerations

  • The project is new and changes quickly. Its recent activity and release cadence are strong, but teams should expect behavior and supported-model details to evolve.
  • Disk capacity and drive speed matter: even the smallest listed model needs substantial storage, and streamed experts can constrain throughput. The README recommends at least 8 GB RAM and more for a better experience.
  • Python 3.10 or newer is required by the project tooling. GPU acceleration is optional, but choosing a GPU backend can involve additional drivers or toolkits.
  • The repository is Apache-2.0 licensed, but model weights retain their own licenses. The README specifically notes that Qwen-Image-2.1 is non-commercial, so check each model's terms before use.
  • The repository's health data does not report a test suite, although the README describes CI checks and a tests directory. Verify the project's checks against your own adoption requirements.

Found this useful?

Share it with someone who would like colibri.

Comparisons