Open Source Inference Projects

Inference is the process of using a trained machine learning model to produce outputs from new inputs. It powers tasks such as generating text, recognizing speech, classifying images, and making predictions. Inference tools help run models efficiently, whether on a personal device, a server, or a distributed system. They address practical needs such as reducing latency and resource use, supporting different hardware, and making model capabilities available through applications and services.

Open source tools in this area include model runtimes, optimization libraries, serving frameworks, and systems for routing requests across models or hardware. When choosing one, consider its maturity, license, maintenance activity, supported models and devices, resource requirements, and integration with your existing software. These tools are useful to developers and researchers building AI applications, as well as teams deploying models locally or at scale.

4 repositories · updated September 25, 2026

llm-d Router: Intelligent Orchestration for Multi-Phase LLM Inference

llm-d Router: Intelligent Orchestration for Multi-Phase LLM Inference

The `llm-d-router` is a sophisticated Go service designed as an intelligent entry point for large language model (LLM) inference requests. It orchestrates complex, multi-phase LLM inference pipelines across specialized worker pools, routing requests through an Inference Gateway to disaggregated vLLM workers. By exposing OpenAI-compatible APIs, it simplifies integration and leverages Kubernetes for scalable, efficient deployments.

AIGateway APIInference
Added Sep 25, 2026 View details
Optimum: Accelerate Hugging Face Models with Hardware Optimization

Optimum: Accelerate Hugging Face Models with Hardware Optimization

Optimum is an extension of Hugging Face Transformers, Diffusers, TIMM, and Sentence-Transformers, designed to provide a suite of optimization tools. It enables maximum efficiency for training and running models on targeted hardware, simplifying the process for developers. This library helps users achieve significant performance gains across various machine learning workflows.

PythonMachine LearningAI
Added Jan 6, 2026 View details
Text Generation Inference: High-Performance LLM Serving by Hugging Face

Text Generation Inference: High-Performance LLM Serving by Hugging Face

Text Generation Inference (TGI) is a robust toolkit from Hugging Face designed for deploying and serving Large Language Models (LLMs) with high performance. It powers Hugging Face's production services, including Hugging Chat and their Inference API. TGI offers optimized text generation, supporting popular open-source LLMs and implementing advanced features for efficient and scalable inference.

Deep LearningInferenceNLP
Added Nov 4, 2025 View details
LitServe: Build Custom Inference Engines for AI Models

LitServe: Build Custom Inference Engines for AI Models

LitServe is a powerful framework from Lightning AI designed to help developers build custom inference engines for a wide range of AI models and systems. It provides expert control over serving, supporting agents, multi-modal systems, RAG, and pipelines without the typical MLOps overhead. This framework offers a flexible and efficient solution for deploying AI models, whether self-hosted or managed on the Lightning AI platform.

PythonAIInference
Added Oct 29, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️