cactus: Run AI Inference on Phones and Wearables

Summary
Cactus is a C++ inference engine for running language, vision, and speech models on mobile and edge devices. It combines quantization, device-focused kernels, and optional cloud handoff for applications that need local inference with a fallback for harder queries.
At a glance
- Language
- C++
- License
- NOASSERTION
- Stars
- 6.1k
- Forks
- 512
- Added to OSRepos
- January 25, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Cactus is a hybrid edge-cloud AI engine for mobile devices and wearables. It combines a model runtime, computation graph, device-oriented kernels, and custom quantization to support local inference for text, vision, and speech.
It is aimed at developers building AI features for phones, tablets, wearables, smart-home devices, and robots. Its local-first approach can reduce dependence on network access, while optional cloud handoff provides a route for queries a local model cannot handle confidently.
Key Features
- Runs language, vision, and speech inference through a shared engine.
- Provides a C API for chat completion, streaming, tool calling, transcription, embeddings, retrieval-augmented generation, and related engine functions.
- Includes a C++ computation graph and CPU/GPU kernels, including ARM NEON kernels.
- Converts model weights with Cactus Quants, using rotation-based quantization options from 1-bit to 4-bit, including mixed-precision modes.
- Offers an OpenAI-compatible local HTTP server and command-line tools for running, converting, downloading, and benchmarking models.
- Supports optional cloud handoff based on local model confidence.
- Provides bindings for Swift, Kotlin, Flutter, React Native, Python, and Rust.
Use Cases
- Mobile app developers can add on-device chat or vision features where network availability or response latency matters.
- Wearable and embedded-device teams can evaluate compact quantized models for constrained hardware.
- Speech application developers can run transcription on supported devices without routing every request to a remote service.
- Teams building assistants can use local tool calling and retrieval, with cloud handoff available for difficult queries.
Project Facts
- Language: C++
- License: NOASSERTION
- Stars: 6.1k
- Forks: 512
- Topics: ai, android, arm, edge, edge-ai, framework, ios, llamacpp, llm, llm-inference, llms, mobile, mobile-inference, on-device-ai, quantiz, rag, smartphone, speech, transformer, whisper
- Archived: false
Getting Started
On a Mac, install and run the CLI:
brew install cactus-compute/cactus/cactus
cactus run
See the repository README for build instructions, model commands, platform options, and API details.
Alternatives
- whisper.cpp: whisper.cpp focuses on offline speech recognition, while Cactus supports language, vision, and speech inference on mobile and edge devices.
- colibri: colibri focuses on streaming Mixture-of-Experts weights from disk, while Cactus targets broader model types and mobile and edge deployment.
Considerations
- Hardware-specific performance varies, and the README's benchmark results cover only the listed devices and models.
- The README describes conversion for arbitrary Hugging Face models as experimental. Liquid, Gemma, Whisper, Parakeet, and Qwen families are identified as especially tested.
- Building from source on Linux requires Python, CMake, build tools, and libcurl development headers, according to the README.
- The repository reports its license as NOASSERTION, so verify licensing terms before adopting it in a project.
- Cloud handoff is optional, but using it may require cloud credentials and introduces a network dependency for those requests.
Source repository
Open the original repository on GitHub.
22 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

SwarmLLM: Run Local and Distributed AI Models
October 2, 2026
SwarmLLM runs open AI models on your computer and can pool resources with other computers to run larger models. It also provides OpenAI- and Anthropic-compatible APIs for local apps and agents.

ai-session-search: Search Local AI Coding Sessions
September 29, 2026
ai-session-search indexes local transcripts from multiple AI coding tools and lets you search sessions, messages, and edited files. Use it from a Rust-powered CLI, MCP server, Rust library, or Python API.

router: Route AI Requests to the Best Model
September 28, 2026
weave-os/router is a Go proxy that routes AI requests across configured model providers, while accepting Anthropic, OpenAI, and Gemini API formats. It suits developers who want model choice and routing behind one endpoint, including agent and coding-tool users.

llm-d-router: Route Inference Requests Intelligently
September 25, 2026
llm-d Router directs inference requests using model-serving signals such as KV-cache locality, load, and priority. It is for teams running LLM serving on Kubernetes that need proxy-integrated routing and request flow control.