LLM Inference

LLM inference is the process of using a trained large language model to produce responses, summaries, code, or other outputs from input prompts. Inference software handles model loading, token generation, and access to computing hardware. It helps teams run language models for applications, manage response speed and resource use, and choose between hosted services and local execution where privacy, connectivity, or cost matters.

Open source tools in this area include inference engines, model servers, API layers, and optimizations for specific hardware or constrained devices. When choosing one, consider model and hardware compatibility, memory and performance requirements, license terms, maintenance activity, and how well it integrates with existing applications. These tools are useful to developers, researchers, and organizations building language model features, from data centers to edge and mobile environments.

3 repositories · updated September 28, 2026

Shimmy: A Pure-Rust WebGPU Inference Engine for GGUF Models

Shimmy: A Pure-Rust WebGPU Inference Engine for GGUF Models

Shimmy is a high-performance, pure-Rust WebGPU inference engine designed for GGUF models. It offers OpenAI-API compatibility, enabling local and private execution of large language models without Python or C++ dependencies. This single-binary solution provides rapid startup and a low memory footprint, making it an efficient alternative for local AI inference.

API ServerLLM InferenceRust
Added Sep 28, 2026 View details
Cactus: Cross-Platform AI Inference Engine for Mobile Devices

Cactus: Cross-Platform AI Inference Engine for Mobile Devices

Cactus is an open-source project providing an energy-efficient, cross-platform AI inference engine specifically designed for mobile devices. It features low-level ARM-specific SIMD operations, a unified zero-copy computation graph, and a high-level transformer engine with NPU support. This framework enables developers to deploy state-of-the-art AI models on smartphones and other edge devices with impressive performance.

AIAndroidEdge
Added Jan 25, 2026 View details
vLLM CLI: A Powerful Command-Line Interface for Serving LLMs with vLLM

vLLM CLI: A Powerful Command-Line Interface for Serving LLMs with vLLM

vLLM CLI is an intuitive command-line interface tool designed to simplify serving Large Language Models using vLLM. It offers both interactive and direct CLI modes, enabling efficient model management, real-time server monitoring, and advanced configuration. This tool streamlines the deployment and management of LLMs, making it accessible for various use cases.

LLMLLM InferenceLLM Tools
Added Jan 20, 2026 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️