Shimmy: A Pure-Rust WebGPU Inference Engine for GGUF Models
This repository profile is provided by osrepos.com, an open source repository discovery platform.

Summary
Shimmy is a high-performance, pure-Rust WebGPU inference engine designed for GGUF models. It offers OpenAI-API compatibility, enabling local and private execution of large language models without Python or C++ dependencies. This single-binary solution provides rapid startup and a low memory footprint, making it an efficient alternative for local AI inference.
Repository Information
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Introduction
Shimmy is an innovative, pure-Rust WebGPU inference engine that serves GGUF models with OpenAI-API compatibility. It stands out by offering a lightweight, single-binary solution for local AI inference, eliminating the need for Python runtimes or C++ toolchains. Built on the Airframe engine, Shimmy provides a robust and efficient platform for running various large language models directly on your GPU.
Why Use Shimmy and Key Advantages
Shimmy offers compelling advantages for developers and users seeking efficient and reliable local LLM inference:
- Pure Rust, Zero Dependencies: Entirely written in Rust, Shimmy requires no Python runtime or C++ toolchain, simplifying deployment and reducing overhead.
- WebGPU Powered: Leverages WGSL compute shaders via WebGPU, ensuring broad compatibility across NVIDIA, AMD, Intel, integrated GPUs, and Apple Silicon.
- OpenAI-API Compatible: Seamlessly integrates with existing OpenAI SDKs and tools, allowing you to point your applications at Shimmy for local, private inference.
- GGUF Native: Directly loads GGUF models, with model specifications automatically derived from metadata, eliminating the need for hardcoded constants.
- Exceptional Performance: Boasts sub-1-second startup times and a minimal memory footprint (around 50MB), significantly outperforming alternatives like Ollama in tested configurations.
- TurboShimmy INT4 KV Cache: Achieves approximately 7x lower KV-cache memory, enabling larger models to run on GPUs with limited VRAM.
- Certified Models: Supports 12 model families and 26 certified model/quant combinations, each passing a rigorous 3-box certification regimen (MATH, INFERENCE, DETERMINISM).
- Extended Context: Features YaRN RoPE scaling for extended context via the
SHIMMY_MAX_CTXenvironment variable. - Deterministic Output: Ensures consistent results, where the same model, seed, and parameters always yield the same output.
Installation
Getting started with Shimmy is straightforward. You can install it using cargo and then serve a model:
cargo install shimmy
shimmy serve --model-path /absolute/path/to/model.gguf --bind 127.0.0.1:11435
For detailed installation instructions, model acquisition, GPU setup, VRAM sizing, and platform-specific builds, please refer to the official Quick Start guide.
Examples
Once Shimmy is running, you can interact with it using simple commands. First, list available models:
shimmy list --short
Then, send a chat completion request:
curl -s http://127.0.0.1:11435/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"tinyllama-1.1b","messages":[{"role":"user","content":"Say hi in 5 words."}],"max_tokens":32}'
Links
- GitHub Repository: https://github.com/Michael-A-Kuykendall/shimmy
- Official Documentation: https://github.com/Michael-A-Kuykendall/shimmy/tree/main/docs
- Quick Start Guide: https://github.com/Michael-A-Kuykendall/shimmy/blob/main/docs/quickstart.md
- Sponsor Shimmy: https://github.com/sponsors/Michael-A-Kuykendall
Source repository
Open the original repository on GitHub.