cactus: Run AI Inference on Phones and Wearables

cactus: Run AI Inference on Phones and Wearables

Summary

Cactus is a C++ inference engine for running language, vision, and speech models on mobile and edge devices. It combines quantization, device-focused kernels, and optional cloud handoff for applications that need local inference with a fallback for harder queries.

At a glance

Language
C++
License
NOASSERTION
Stars
6.1k
Forks
512
Added to OSRepos
January 25, 2026
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

Cactus is a hybrid edge-cloud AI engine for mobile devices and wearables. It combines a model runtime, computation graph, device-oriented kernels, and custom quantization to support local inference for text, vision, and speech.

It is aimed at developers building AI features for phones, tablets, wearables, smart-home devices, and robots. Its local-first approach can reduce dependence on network access, while optional cloud handoff provides a route for queries a local model cannot handle confidently.

Key Features

  • Runs language, vision, and speech inference through a shared engine.
  • Provides a C API for chat completion, streaming, tool calling, transcription, embeddings, retrieval-augmented generation, and related engine functions.
  • Includes a C++ computation graph and CPU/GPU kernels, including ARM NEON kernels.
  • Converts model weights with Cactus Quants, using rotation-based quantization options from 1-bit to 4-bit, including mixed-precision modes.
  • Offers an OpenAI-compatible local HTTP server and command-line tools for running, converting, downloading, and benchmarking models.
  • Supports optional cloud handoff based on local model confidence.
  • Provides bindings for Swift, Kotlin, Flutter, React Native, Python, and Rust.

Use Cases

  • Mobile app developers can add on-device chat or vision features where network availability or response latency matters.
  • Wearable and embedded-device teams can evaluate compact quantized models for constrained hardware.
  • Speech application developers can run transcription on supported devices without routing every request to a remote service.
  • Teams building assistants can use local tool calling and retrieval, with cloud handoff available for difficult queries.

Project Facts

  • Language: C++
  • License: NOASSERTION
  • Stars: 6.1k
  • Forks: 512
  • Topics: ai, android, arm, edge, edge-ai, framework, ios, llamacpp, llm, llm-inference, llms, mobile, mobile-inference, on-device-ai, quantiz, rag, smartphone, speech, transformer, whisper
  • Archived: false

Getting Started

On a Mac, install and run the CLI:

brew install cactus-compute/cactus/cactus
cactus run

See the repository README for build instructions, model commands, platform options, and API details.

Alternatives

  • whisper.cpp: whisper.cpp focuses on offline speech recognition, while Cactus supports language, vision, and speech inference on mobile and edge devices.
  • colibri: colibri focuses on streaming Mixture-of-Experts weights from disk, while Cactus targets broader model types and mobile and edge deployment.

Considerations

  • Hardware-specific performance varies, and the README's benchmark results cover only the listed devices and models.
  • The README describes conversion for arbitrary Hugging Face models as experimental. Liquid, Gemma, Whisper, Parakeet, and Qwen families are identified as especially tested.
  • Building from source on Linux requires Python, CMake, build tools, and libcurl development headers, according to the README.
  • The repository reports its license as NOASSERTION, so verify licensing terms before adopting it in a project.
  • Cloud handoff is optional, but using it may require cloud credentials and introduces a network dependency for those requests.

Source repository

Open the original repository on GitHub.

22 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design

web-design: A Claude Code SKILL for Spec-First Web Page Design

October 3, 2026

The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

Claude CodeClaude SkillDesign System
OrcaReplay: Time Travel for AI Agents, Debugging and Evaluation

OrcaReplay: Time Travel for AI Agents, Debugging and Evaluation

October 2, 2026

OrcaReplay introduces "time travel" capabilities for AI agents, allowing developers to record, replay, fork, and debug any agent run with any model. It addresses the challenges of AI agent debugging by providing byte-for-byte reproducibility, offline analysis, and the ability to compare different models from specific checkpoints. This tool, built by the OrcaRouter.ai team, enhances observability and control over complex agent behaviors.

Agent DebuggingAI AgentsLLM Agents
Shepherd: Reversible Execution Traces for Programmable Meta-Agents

Shepherd: Reversible Execution Traces for Programmable Meta-Agents

October 2, 2026

Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

PythonAIAgent Framework
SwarmLLM: Run Local AI Models and Team Up for Giant Distributed Inference

SwarmLLM: Run Local AI Models and Team Up for Giant Distributed Inference

October 2, 2026

SwarmLLM is a free, open-source application that allows you to run AI chat models directly on your own computer. It uniquely enables multiple computers to team up over the internet, collectively running models too large for a single machine. This platform offers an OpenAI and Anthropic-compatible API, all without requiring accounts or cryptocurrency.

AILLMDecentralized
OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️