{"name":"Shimmy: A Pure-Rust WebGPU Inference Engine for GGUF Models","description":"Shimmy is a high-performance, pure-Rust WebGPU inference engine designed for GGUF models. It offers OpenAI-API compatibility, enabling local and private execution of large language models without Python or C++ dependencies. This single-binary solution provides rapid startup and a low memory footprint, making it an efficient alternative for local AI inference.","github":"https://github.com/Michael-A-Kuykendall/shimmy","url":"https://osrepos.com/repo/michael-a-kuykendall-shimmy","source":"osrepos.com","sourceDescription":"This repository profile is provided by osrepos.com, an open source repository discovery platform.","repositoryProfile":"https://osrepos.com/repo/michael-a-kuykendall-shimmy","generatedFor":"open source discovery and AI-assisted research","markdown":"https://osrepos.com/repo/michael-a-kuykendall-shimmy.md","json":"https://osrepos.com/repo/michael-a-kuykendall-shimmy.json","topics":["api-server","llm-inference","rust","webgpu","machine-learning","developer-tools","openai-compatible","gguf"],"keywords":["api-server","llm-inference","rust","webgpu","machine-learning","developer-tools","openai-compatible","gguf"],"stars":null,"summary":"Shimmy is a high-performance, pure-Rust WebGPU inference engine designed for GGUF models. It offers OpenAI-API compatibility, enabling local and private execution of large language models without Python or C++ dependencies. This single-binary solution provides rapid startup and a low memory footprint, making it an efficient alternative for local AI inference.","content":"## Introduction\nShimmy is an innovative, pure-Rust WebGPU inference engine that serves GGUF models with OpenAI-API compatibility. It stands out by offering a lightweight, single-binary solution for local AI inference, eliminating the need for Python runtimes or C++ toolchains. Built on the Airframe engine, Shimmy provides a robust and efficient platform for running various large language models directly on your GPU.\n\n## Why Use Shimmy and Key Advantages\nShimmy offers compelling advantages for developers and users seeking efficient and reliable local LLM inference:\n*   **Pure Rust, Zero Dependencies**: Entirely written in Rust, Shimmy requires no Python runtime or C++ toolchain, simplifying deployment and reducing overhead.\n*   **WebGPU Powered**: Leverages WGSL compute shaders via WebGPU, ensuring broad compatibility across NVIDIA, AMD, Intel, integrated GPUs, and Apple Silicon.\n*   **OpenAI-API Compatible**: Seamlessly integrates with existing OpenAI SDKs and tools, allowing you to point your applications at Shimmy for local, private inference.\n*   **GGUF Native**: Directly loads GGUF models, with model specifications automatically derived from metadata, eliminating the need for hardcoded constants.\n*   **Exceptional Performance**: Boasts sub-1-second startup times and a minimal memory footprint (around 50MB), significantly outperforming alternatives like Ollama in tested configurations.\n*   **TurboShimmy INT4 KV Cache**: Achieves approximately 7x lower KV-cache memory, enabling larger models to run on GPUs with limited VRAM.\n*   **Certified Models**: Supports 12 model families and 26 certified model/quant combinations, each passing a rigorous 3-box certification regimen (MATH, INFERENCE, DETERMINISM).\n*   **Extended Context**: Features YaRN RoPE scaling for extended context via the `SHIMMY_MAX_CTX` environment variable.\n*   **Deterministic Output**: Ensures consistent results, where the same model, seed, and parameters always yield the same output.\n\n## Installation\nGetting started with Shimmy is straightforward. You can install it using `cargo` and then serve a model:\n\nbash\ncargo install shimmy\nshimmy serve --model-path /absolute/path/to/model.gguf --bind 127.0.0.1:11435\n\n\nFor detailed installation instructions, model acquisition, GPU setup, VRAM sizing, and platform-specific builds, please refer to the official [Quick Start guide](https://github.com/Michael-A-Kuykendall/shimmy/blob/main/docs/quickstart.md).\n\n## Examples\nOnce Shimmy is running, you can interact with it using simple commands. First, list available models:\n\nbash\nshimmy list --short\n\n\nThen, send a chat completion request:\n\nbash\ncurl -s http://127.0.0.1:11435/v1/chat/completions \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"model\":\"tinyllama-1.1b\",\"messages\":[{\"role\":\"user\",\"content\":\"Say hi in 5 words.\"}],\"max_tokens\":32}'\n\n\n## Links\n*   **GitHub Repository**: [https://github.com/Michael-A-Kuykendall/shimmy](https://github.com/Michael-A-Kuykendall/shimmy)\n*   **Official Documentation**: [https://github.com/Michael-A-Kuykendall/shimmy/tree/main/docs](https://github.com/Michael-A-Kuykendall/shimmy/tree/main/docs)\n*   **Quick Start Guide**: [https://github.com/Michael-A-Kuykendall/shimmy/blob/main/docs/quickstart.md](https://github.com/Michael-A-Kuykendall/shimmy/blob/main/docs/quickstart.md)\n*   **Sponsor Shimmy**: [https://github.com/sponsors/Michael-A-Kuykendall](https://github.com/sponsors/Michael-A-Kuykendall)","metrics":{"detailViews":0,"githubClicks":0},"dates":{"published":null,"modified":"2026-09-28T11:32:18.000Z"}}