infinity: Serve Embedding and Reranking Models via API

Summary
Infinity is a Python serving engine for text embeddings, reranking, and selected image, audio, and late-interaction models. It provides an OpenAI-aligned REST API with multiple inference backends for teams deploying models from Hugging Face.
At a glance
- Language
- Python
- License
- MIT
- Stars
- 2.9k
- Forks
- 209
- Added to OSRepos
- March 17, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Infinity runs embedding and reranking models behind a REST API, addressing the need to host these inference workloads rather than rely on a hosted provider. It also supports selected multimodal and late-interaction tasks, including CLIP, CLAP, ColBERT-style, and ColPali-style models.
The project is aimed at teams building search, retrieval-augmented generation, or other applications that need model inference as a service. Its API is aligned with OpenAI's embeddings specification, and it can serve multiple models from one process.
Key Features
- Serves text embedding, reranking, and text classification models.
- Supports selected image-text, audio-text, ColBERT-style, and ColPali-style embedding tasks.
- Offers PyTorch, Optimum (ONNX/TensorRT), and CTranslate2 inference engines, with hardware options depending on backend and setup.
- Uses dynamic batching and worker threads for tokenization.
- Exposes a FastAPI REST API and an asynchronous Python API.
- Provides CLI and environment-variable configuration, including launching multiple models.
- Offers Docker images for standard and specialized deployments.
Use Cases
- Search and recommendation teams can generate embeddings from self-hosted models for indexing and query processing.
- RAG developers can deploy embedding and reranking endpoints alongside their application and vector database.
- Platform teams can offer shared inference endpoints for multiple models through one service.
- Applications working with image or audio search can serve supported CLIP- or CLAP-based models.
- Python developers can call model engines directly through the asynchronous API instead of making REST requests.
Project Facts
- Language: Python
- License: MIT
- Stars: 2.9k
- Forks: 209
- Topics: bert-embeddings, llm, text-embeddings
- Archived: false
Getting Started
Install the package and launch a model:
pip install infinity-emb[all]
infinity_emb v2 --model-id BAAI/bge-small-en-v1.5
See the README and documentation for Docker, backend, hardware, and API configuration details.
Considerations
- Model compatibility depends on the selected inference engine. For example, the README describes CTranslate2 support as limited to BERT models and Optimum use as requiring an ONNX model.
- Hardware acceleration is not automatic. GPU deployments require compatible drivers, libraries, and container configuration; supported options vary by backend.
- The repository describes some multimodal and late-interaction support, but lists specific tested model families and also notes unsupported model types.
- Some specialized Docker images are not built through CI/CD, and the README recommends pinning an exact version for those images.
- Serving models locally means operators must provision compute, manage model files and caching, and secure the API as appropriate for their deployment.
Source repository
Open the original repository on GitHub.
20 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design
October 3, 2026
The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner
October 2, 2026
OOMWOO is an ambitious open-source project enabling users to build their own robot vacuum cleaner using Raspberry Pi, 3D printing, and ROS2. It emphasizes local operation, hackability, and integration with Home Assistant, providing a high-quality, customizable home appliance. This project aims to deliver a fully open hardware, software, and firmware solution for autonomous home cleaning.

Shepherd: Reversible Execution Traces for Programmable Meta-Agents
October 2, 2026
Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents
October 1, 2026
Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.