cactus: Run AI Inference on Phones and Wearables

Summary
Cactus is a C++ inference engine for running language, vision, and speech models on mobile and edge devices. It combines quantization, device-focused kernels, and optional cloud handoff for applications that need local inference with a fallback for harder queries.
At a glance
- Language
- C++
- License
- NOASSERTION
- Stars
- 6.1k
- Forks
- 512
- Added to OSRepos
- January 25, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Cactus is a hybrid edge-cloud AI engine for mobile devices and wearables. It combines a model runtime, computation graph, device-oriented kernels, and custom quantization to support local inference for text, vision, and speech.
It is aimed at developers building AI features for phones, tablets, wearables, smart-home devices, and robots. Its local-first approach can reduce dependence on network access, while optional cloud handoff provides a route for queries a local model cannot handle confidently.
Key Features
- Runs language, vision, and speech inference through a shared engine.
- Provides a C API for chat completion, streaming, tool calling, transcription, embeddings, retrieval-augmented generation, and related engine functions.
- Includes a C++ computation graph and CPU/GPU kernels, including ARM NEON kernels.
- Converts model weights with Cactus Quants, using rotation-based quantization options from 1-bit to 4-bit, including mixed-precision modes.
- Offers an OpenAI-compatible local HTTP server and command-line tools for running, converting, downloading, and benchmarking models.
- Supports optional cloud handoff based on local model confidence.
- Provides bindings for Swift, Kotlin, Flutter, React Native, Python, and Rust.
Use Cases
- Mobile app developers can add on-device chat or vision features where network availability or response latency matters.
- Wearable and embedded-device teams can evaluate compact quantized models for constrained hardware.
- Speech application developers can run transcription on supported devices without routing every request to a remote service.
- Teams building assistants can use local tool calling and retrieval, with cloud handoff available for difficult queries.
Project Facts
- Language: C++
- License: NOASSERTION
- Stars: 6.1k
- Forks: 512
- Topics: ai, android, arm, edge, edge-ai, framework, ios, llamacpp, llm, llm-inference, llms, mobile, mobile-inference, on-device-ai, quantiz, rag, smartphone, speech, transformer, whisper
- Archived: false
Getting Started
On a Mac, install and run the CLI:
brew install cactus-compute/cactus/cactus
cactus run
See the repository README for build instructions, model commands, platform options, and API details.
Alternatives
- whisper.cpp: whisper.cpp focuses on offline speech recognition, while Cactus supports language, vision, and speech inference on mobile and edge devices.
- colibri: colibri focuses on streaming Mixture-of-Experts weights from disk, while Cactus targets broader model types and mobile and edge deployment.
Considerations
- Hardware-specific performance varies, and the README's benchmark results cover only the listed devices and models.
- The README describes conversion for arbitrary Hugging Face models as experimental. Liquid, Gemma, Whisper, Parakeet, and Qwen families are identified as especially tested.
- Building from source on Linux requires Python, CMake, build tools, and libcurl development headers, according to the README.
- The repository reports its license as NOASSERTION, so verify licensing terms before adopting it in a project.
- Cloud handoff is optional, but using it may require cloud credentials and introduces a network dependency for those requests.
Source repository
Open the original repository on GitHub.
22 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design
October 3, 2026
The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

OrcaReplay: Time Travel for AI Agents, Debugging and Evaluation
October 2, 2026
OrcaReplay introduces "time travel" capabilities for AI agents, allowing developers to record, replay, fork, and debug any agent run with any model. It addresses the challenges of AI agent debugging by providing byte-for-byte reproducibility, offline analysis, and the ability to compare different models from specific checkpoints. This tool, built by the OrcaRouter.ai team, enhances observability and control over complex agent behaviors.

Shepherd: Reversible Execution Traces for Programmable Meta-Agents
October 2, 2026
Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

SwarmLLM: Run Local AI Models and Team Up for Giant Distributed Inference
October 2, 2026
SwarmLLM is a free, open-source application that allows you to run AI chat models directly on your own computer. It uniquely enables multiple computers to team up over the internet, collectively running models too large for a single machine. This platform offers an OpenAI and Anthropic-compatible API, all without requiring accounts or cryptocurrency.