Shimmy: A Pure-Rust WebGPU Inference Engine for GGUF Models

This repository profile is provided by osrepos.com, an open source repository discovery platform.

Shimmy: A Pure-Rust WebGPU Inference Engine for GGUF Models

Summary

Shimmy is a high-performance, pure-Rust WebGPU inference engine designed for GGUF models. It offers OpenAI-API compatibility, enabling local and private execution of large language models without Python or C++ dependencies. This single-binary solution provides rapid startup and a low memory footprint, making it an efficient alternative for local AI inference.

Repository Information

Analyzed by OSRepos on September 28, 2026

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Introduction

Shimmy is an innovative, pure-Rust WebGPU inference engine that serves GGUF models with OpenAI-API compatibility. It stands out by offering a lightweight, single-binary solution for local AI inference, eliminating the need for Python runtimes or C++ toolchains. Built on the Airframe engine, Shimmy provides a robust and efficient platform for running various large language models directly on your GPU.

Why Use Shimmy and Key Advantages

Shimmy offers compelling advantages for developers and users seeking efficient and reliable local LLM inference:

  • Pure Rust, Zero Dependencies: Entirely written in Rust, Shimmy requires no Python runtime or C++ toolchain, simplifying deployment and reducing overhead.
  • WebGPU Powered: Leverages WGSL compute shaders via WebGPU, ensuring broad compatibility across NVIDIA, AMD, Intel, integrated GPUs, and Apple Silicon.
  • OpenAI-API Compatible: Seamlessly integrates with existing OpenAI SDKs and tools, allowing you to point your applications at Shimmy for local, private inference.
  • GGUF Native: Directly loads GGUF models, with model specifications automatically derived from metadata, eliminating the need for hardcoded constants.
  • Exceptional Performance: Boasts sub-1-second startup times and a minimal memory footprint (around 50MB), significantly outperforming alternatives like Ollama in tested configurations.
  • TurboShimmy INT4 KV Cache: Achieves approximately 7x lower KV-cache memory, enabling larger models to run on GPUs with limited VRAM.
  • Certified Models: Supports 12 model families and 26 certified model/quant combinations, each passing a rigorous 3-box certification regimen (MATH, INFERENCE, DETERMINISM).
  • Extended Context: Features YaRN RoPE scaling for extended context via the SHIMMY_MAX_CTX environment variable.
  • Deterministic Output: Ensures consistent results, where the same model, seed, and parameters always yield the same output.

Installation

Getting started with Shimmy is straightforward. You can install it using cargo and then serve a model:

cargo install shimmy
shimmy serve --model-path /absolute/path/to/model.gguf --bind 127.0.0.1:11435

For detailed installation instructions, model acquisition, GPU setup, VRAM sizing, and platform-specific builds, please refer to the official Quick Start guide.

Examples

Once Shimmy is running, you can interact with it using simple commands. First, list available models:

shimmy list --short

Then, send a chat completion request:

curl -s http://127.0.0.1:11435/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"tinyllama-1.1b","messages":[{"role":"user","content":"Say hi in 5 words."}],"max_tokens":32}'

Links

Source repository

Open the original repository on GitHub.

View on GitHub
OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️