Open Source Inference Projects
Inference is the process of using a trained machine learning model to produce outputs from new inputs. It powers tasks such as generating text, recognizing speech, classifying images, and making predictions. Inference tools help run models efficiently, whether on a personal device, a server, or a distributed system. They address practical needs such as reducing latency and resource use, supporting different hardware, and making model capabilities available through applications and services.
Open source tools in this area include model runtimes, optimization libraries, serving frameworks, and systems for routing requests across models or hardware. When choosing one, consider its maturity, license, maintenance activity, supported models and devices, resource requirements, and integration with your existing software. These tools are useful to developers and researchers building AI applications, as well as teams deploying models locally or at scale.
4 repositories · updated September 25, 2026

llm-d Router: Intelligent Orchestration for Multi-Phase LLM Inference
The `llm-d-router` is a sophisticated Go service designed as an intelligent entry point for large language model (LLM) inference requests. It orchestrates complex, multi-phase LLM inference pipelines across specialized worker pools, routing requests through an Inference Gateway to disaggregated vLLM workers. By exposing OpenAI-compatible APIs, it simplifies integration and leverages Kubernetes for scalable, efficient deployments.

Optimum: Accelerate Hugging Face Models with Hardware Optimization
Optimum is an extension of Hugging Face Transformers, Diffusers, TIMM, and Sentence-Transformers, designed to provide a suite of optimization tools. It enables maximum efficiency for training and running models on targeted hardware, simplifying the process for developers. This library helps users achieve significant performance gains across various machine learning workflows.

Text Generation Inference: High-Performance LLM Serving by Hugging Face
Text Generation Inference (TGI) is a robust toolkit from Hugging Face designed for deploying and serving Large Language Models (LLMs) with high performance. It powers Hugging Face's production services, including Hugging Chat and their Inference API. TGI offers optimized text generation, supporting popular open-source LLMs and implementing advanced features for efficient and scalable inference.

LitServe: Build Custom Inference Engines for AI Models
LitServe is a powerful framework from Lightning AI designed to help developers build custom inference engines for a wide range of AI models and systems. It provides expert control over serving, supporting agents, multi-modal systems, RAG, and pipelines without the typical MLOps overhead. This framework offers a flexible and efficient solution for deploying AI models, whether self-hosted or managed on the Lightning AI platform.