Repository History
3 repositories tagged with llm-inference

Shimmy: A Pure-Rust WebGPU Inference Engine for GGUF Models
Shimmy is a high-performance, pure-Rust WebGPU inference engine designed for GGUF models. It offers OpenAI-API compatibility, enabling local and private execution of large language models without Python or C++ dependencies. This single-binary solution provides rapid startup and a low memory footprint, making it an efficient alternative for local AI inference.

Cactus: Cross-Platform AI Inference Engine for Mobile Devices
Cactus is an open-source project providing an energy-efficient, cross-platform AI inference engine specifically designed for mobile devices. It features low-level ARM-specific SIMD operations, a unified zero-copy computation graph, and a high-level transformer engine with NPU support. This framework enables developers to deploy state-of-the-art AI models on smartphones and other edge devices with impressive performance.

vLLM CLI: A Powerful Command-Line Interface for Serving LLMs with vLLM
vLLM CLI is an intuitive command-line interface tool designed to simplify serving Large Language Models using vLLM. It offers both interactive and direct CLI modes, enabling efficient model management, real-time server monitoring, and advanced configuration. This tool streamlines the deployment and management of LLMs, making it accessible for various use cases.