{"name":"Cactus: Cross-Platform AI Inference Engine for Mobile Devices","description":"Cactus is an open-source project providing an energy-efficient, cross-platform AI inference engine specifically designed for mobile devices. It features low-level ARM-specific SIMD operations, a unified zero-copy computation graph, and a high-level transformer engine with NPU support. This framework enables developers to deploy state-of-the-art AI models on smartphones and other edge devices with impressive performance.","github":"https://github.com/cactus-compute/cactus","url":"https://osrepos.com/repo/cactus-compute-cactus","source":"osrepos.com","sourceDescription":"This repository profile is provided by osrepos.com, an open source repository discovery platform.","repositoryProfile":"https://osrepos.com/repo/cactus-compute-cactus","generatedFor":"open source discovery and AI-assisted research","markdown":"https://osrepos.com/repo/cactus-compute-cactus.md","json":"https://osrepos.com/repo/cactus-compute-cactus.json","topics":["ai","android","edge","ios","llm-inference","mobile","C++","framework"],"keywords":["ai","android","edge","ios","llm-inference","mobile","C++","framework"],"stars":null,"summary":"Cactus is an open-source project providing an energy-efficient, cross-platform AI inference engine specifically designed for mobile devices. It features low-level ARM-specific SIMD operations, a unified zero-copy computation graph, and a high-level transformer engine with NPU support. This framework enables developers to deploy state-of-the-art AI models on smartphones and other edge devices with impressive performance.","content":"## Introduction\n\nCactus is a powerful open-source project by Cactus Compute, Inc. and UCLA's BruinAI, offering an advanced AI inference engine optimized for mobile devices. It provides a comprehensive stack, from low-level hardware-specific kernels to a high-level API, enabling efficient execution of AI models on various platforms like iOS and Android.\n\nThe architecture of Cactus is layered, starting with `Cactus Kernels` for ARM-specific SIMD operations, building up to `Cactus Graph` for unified computation, `Cactus Models` for implementing SOTA models, `Cactus Engine` for transformer inference with NPU support, and finally `Cactus FFI` for an OpenAI-compatible C API.\n\n## Installation\n\nGetting started with Cactus on a Mac is straightforward. First, clone the repository and navigate into its directory. Then, source the setup script to configure your environment.\n\nbash\ngit clone https://github.com/cactus-compute/cactus && cd cactus && source ./setup\n\n\nRemember to run `source ./setup` in any new terminal session. You can use `cactus run [model]` to open a playground, `cactus download [model]` to get models, and `cactus build` to compile for ARM targets like Apple or Android.\n\n## Examples\n\nCactus provides intuitive APIs for both graph-level operations and high-level model inference. Below are examples demonstrating how to use `Cactus Graph & Kernel` for custom computations and `Cactus Engine & FFI` for running LLM inference.\n\n### Cactus Graph & Kernel\n\nThis example shows how to define a computation graph, set inputs, execute it, and retrieve results using Cactus's low-level graph API.\n\ncpp\n#include cactus.h\n\nCactusGraph graph;\nauto a = graph.input({2, 3}, Precision::FP16);\nauto b = graph.input({3, 4}, Precision::INT8);\n\nauto x1 = graph.matmul(a, b, false);\nauto x2 = graph.transpose(x1);\nauto result = graph.matmul(b, x2, true);\n\nfloat a_data[6] = {1.1f, 2.3f, 3.4f, 4.2f, 5.7f, 6.8f};\nfloat b_data[12] = {1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12};\n\ngraph.set_input(a, a_data, Precision::FP16);\ngraph.set_input(b, b_data, Precision::INT8);\n\ngraph.execute();\nvoid* output_data = graph.get_output(result);\n\ngraph.hard_reset(); \n\n\n\n### Cactus Engine & FFI\n\nThis snippet demonstrates using the `Cactus Engine` via its FFI to perform AI inference, similar to OpenAI's API, with support for chat messages and generation options.\n\ncpp\n#include cactus.h\n\ncactus_set_pro_key(\"\");                  // email founders@cactuscompute.com for optional key\n\ncactus_model_t model = cactus_init(\n    \"path/to/weight/folder\",             // section to generate weigths below\n    \"txt/or/md/file/or/dir/with/many\",   // nullptr if none, cactus does automatic fast RAG\n);\n\nconst char* messages = R\"([\n    {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n    {\"role\": \"user\", \"content\": \"My name is Henry Ndubuaku\"}\n])\";\n\nconst char* options = R\"({\n    \"max_tokens\": 50,\n    \"stop_sequences\": [\"<|im_end|>\"]\n})\";\n\nchar response[4096];\nint result = cactus_complete(\n    model,                               // model handle from cactus_init\n    messages,                            // JSON array of chat messages\n    response,                            // buffer to store response JSON\n    sizeof(response),                    // size of response buffer\n    options,                             // optional: generation options (nullptr for defaults)\n    nullptr,                             // optional: tools JSON for function calling \n    nullptr,                             // optional: streaming callback fn(token, id, user_data)\n    nullptr                              // optional: user data passed to callback\n);\n\n\n## Why Use Cactus\n\nCactus stands out as a robust solution for on-device AI inference due to several key advantages:\n\n*   **Cross-Platform & Energy-Efficient**: Designed from the ground up for mobile devices, ensuring optimal performance and low power consumption on both iOS and Android.\n*   **NPU Support**: Leverages Neural Processing Units (NPUs) on compatible devices, such as Apple Silicon, for real-time inference and handling large contexts.\n*   **Mixed Precision**: Intelligently blends INT4, INT8, and FP16 for weights, achieving optimal balance between model size, speed, and accuracy.\n*   **High Performance**: Offers impressive decode and prefill tokens per second, as demonstrated across various devices from Mac M4 Pro to Raspberry Pi 5.\n*   **Extensive Model Support**: Supports a wide range of popular models including Gemma, Whisper, LFM2, and Qwen, with capabilities for completion, tools, vision, embeddings, and speech.\n*   **Developer-Friendly Ecosystem**: Provides SDKs for Python, React Native, Swift, Kotlin, Flutter, and Rust, making integration into existing applications seamless.\n\n## Links\n\nExplore Cactus further through its official repository and various SDKs and demo applications:\n\n*   **GitHub Repository**: [https://github.com/cactus-compute/cactus](https://github.com/cactus-compute/cactus){:target=\"_blank\"}\n*   **Python SDK for Mac**: [https://github.com/cactus-compute/cactus/tree/main/python](https://github.com/cactus-compute/cactus/tree/main/python){:target=\"_blank\"}\n*   **React Native SDK**: [https://github.com/cactus-compute/cactus-react-native](https://github.com/cactus-compute/cactus-react-native){:target=\"_blank\"}\n*   **Swift Multiplatform SDK**: [https://github.com/mhayes853/swift-cactus](https://github.com/mhayes853/swift-cactus){:target=\"_blank\"}\n*   **Kotlin Multiplatform SDK**: [https://github.com/cactus-compute/cactus-kotlin](https://github.com/cactus-compute/cactus-kotlin){:target=\"_blank\"}\n*   **Flutter SDK**: [https://github.com/cactus-compute/cactus-flutter](https://github.com/cactus-compute/cactus-flutter){:target=\"_blank\"}\n*   **Rust SDK**: [https://github.com/mrsarac/cactus-rs](https://github.com/mrsarac/cactus-rs){:target=\"_blank\"}\n*   **iOS Demo App**: [https://apps.apple.com/gb/app/cactus-chat/id6744444212](https://apps.apple.com/gb/app/cactus-chat/id6744444212){:target=\"_blank\"}\n*   **Android Demo App**: [https://play.google.com/store/apps/details?id=com.rshemetsubuser.myapp](https://play.google.com/store/apps/details?id=com.rshemetsubuser.myapp){:target=\"_blank\"}","metrics":{"detailViews":15,"githubClicks":16},"dates":{"published":null,"modified":"2026-01-25T20:01:08.000Z"}}