# Shimmy: A Pure-Rust WebGPU Inference Engine for GGUF Models

This repository profile is provided by osrepos.com, an open source repository discovery platform.

Source: osrepos.com
Repository profile: https://osrepos.com/repo/michael-a-kuykendall-shimmy
Generated for open source discovery and AI-assisted research.

Shimmy is a high-performance, pure-Rust WebGPU inference engine designed for GGUF models. It offers OpenAI-API compatibility, enabling local and private execution of large language models without Python or C++ dependencies. This single-binary solution provides rapid startup and a low memory footprint, making it an efficient alternative for local AI inference.

GitHub: https://github.com/Michael-A-Kuykendall/shimmy
OSRepos URL: https://osrepos.com/repo/michael-a-kuykendall-shimmy

## Summary

Shimmy is a high-performance, pure-Rust WebGPU inference engine designed for GGUF models. It offers OpenAI-API compatibility, enabling local and private execution of large language models without Python or C++ dependencies. This single-binary solution provides rapid startup and a low memory footprint, making it an efficient alternative for local AI inference.

## Topics

- api-server
- llm-inference
- rust
- webgpu
- machine-learning
- developer-tools
- openai-compatible
- gguf

## Repository Information

Last analyzed by OSRepos: Mon Sep 28 2026 12:32:18 GMT+0100 (Western European Summer Time)
Detail views: 0
GitHub clicks: 0

## Safety Notice

OSRepos shares public repositories for knowledge and discovery only. Review source code, dependencies, licenses, and security implications before running or installing anything.

## Content

## Introduction
Shimmy is an innovative, pure-Rust WebGPU inference engine that serves GGUF models with OpenAI-API compatibility. It stands out by offering a lightweight, single-binary solution for local AI inference, eliminating the need for Python runtimes or C++ toolchains. Built on the Airframe engine, Shimmy provides a robust and efficient platform for running various large language models directly on your GPU.

## Why Use Shimmy and Key Advantages
Shimmy offers compelling advantages for developers and users seeking efficient and reliable local LLM inference:
*   **Pure Rust, Zero Dependencies**: Entirely written in Rust, Shimmy requires no Python runtime or C++ toolchain, simplifying deployment and reducing overhead.
*   **WebGPU Powered**: Leverages WGSL compute shaders via WebGPU, ensuring broad compatibility across NVIDIA, AMD, Intel, integrated GPUs, and Apple Silicon.
*   **OpenAI-API Compatible**: Seamlessly integrates with existing OpenAI SDKs and tools, allowing you to point your applications at Shimmy for local, private inference.
*   **GGUF Native**: Directly loads GGUF models, with model specifications automatically derived from metadata, eliminating the need for hardcoded constants.
*   **Exceptional Performance**: Boasts sub-1-second startup times and a minimal memory footprint (around 50MB), significantly outperforming alternatives like Ollama in tested configurations.
*   **TurboShimmy INT4 KV Cache**: Achieves approximately 7x lower KV-cache memory, enabling larger models to run on GPUs with limited VRAM.
*   **Certified Models**: Supports 12 model families and 26 certified model/quant combinations, each passing a rigorous 3-box certification regimen (MATH, INFERENCE, DETERMINISM).
*   **Extended Context**: Features YaRN RoPE scaling for extended context via the `SHIMMY_MAX_CTX` environment variable.
*   **Deterministic Output**: Ensures consistent results, where the same model, seed, and parameters always yield the same output.

## Installation
Getting started with Shimmy is straightforward. You can install it using `cargo` and then serve a model:

bash
cargo install shimmy
shimmy serve --model-path /absolute/path/to/model.gguf --bind 127.0.0.1:11435


For detailed installation instructions, model acquisition, GPU setup, VRAM sizing, and platform-specific builds, please refer to the official [Quick Start guide](https://github.com/Michael-A-Kuykendall/shimmy/blob/main/docs/quickstart.md).

## Examples
Once Shimmy is running, you can interact with it using simple commands. First, list available models:

bash
shimmy list --short


Then, send a chat completion request:

bash
curl -s http://127.0.0.1:11435/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"tinyllama-1.1b","messages":[{"role":"user","content":"Say hi in 5 words."}],"max_tokens":32}'


## Links
*   **GitHub Repository**: [https://github.com/Michael-A-Kuykendall/shimmy](https://github.com/Michael-A-Kuykendall/shimmy)
*   **Official Documentation**: [https://github.com/Michael-A-Kuykendall/shimmy/tree/main/docs](https://github.com/Michael-A-Kuykendall/shimmy/tree/main/docs)
*   **Quick Start Guide**: [https://github.com/Michael-A-Kuykendall/shimmy/blob/main/docs/quickstart.md](https://github.com/Michael-A-Kuykendall/shimmy/blob/main/docs/quickstart.md)
*   **Sponsor Shimmy**: [https://github.com/sponsors/Michael-A-Kuykendall](https://github.com/sponsors/Michael-A-Kuykendall)