optimum: Optimize Model Training and Inference on Target Hardware

Summary
Hugging Face Optimum adds tools for optimizing model training and inference across hardware backends. It suits teams using Transformers, Diffusers, TIMM, or Sentence Transformers who need hardware-specific deployment or training workflows.
At a glance
- Language
- Python
- License
- Apache-2.0
- Stars
- 3.5k
- Forks
- 695
- Added to OSRepos
- January 6, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Optimum is a Python library that connects Hugging Face model workflows with optimization and acceleration tools for different hardware. It helps developers export, run, and train models using backends such as OpenVINO, ONNX Runtime, ExecuTorch, and accelerator-specific integrations.
It is most useful when a standard Transformers workflow needs to target a particular device or inference runtime. The core package is a starting point, while many backends require separate extras or companion projects.
Key Features
- Integrates optimization workflows with Transformers, Diffusers, TIMM, and Sentence Transformers.
- Supports model export and optimized inference across several runtimes and hardware ecosystems.
- Provides integrations for Intel OpenVINO, Intel Gaudi, AWS Inferentia and Trainium, AMD hardware, NVIDIA TensorRT-LLM, and other providers.
- Offers training integrations for supported accelerators, including Gaudi and AWS Trainium.
- Includes command-line and programmatic workflows for supported export and optimization tasks.
- Connects users to separate projects for some capabilities, including Optimum ONNX, Optimum Quanto, and Optimum ExecuTorch.
Use Cases
- A machine-learning engineer can export a Transformers model for an inference runtime supported by their deployment hardware.
- An application team can evaluate OpenVINO or another supported backend when deploying models on a specific accelerator ecosystem.
- A model trainer can use the Gaudi or Trainium integrations when training on supported hardware.
- An edge developer can use the Optimum ExecuTorch project to export Transformers models for PyTorch-based on-device inference.
Project Facts
- Language: Python
- License: Apache-2.0
- Stars: 3.5k
- Forks: 695
- Topics: graphcore, habana, inference, intel, onnx, onnxruntime, optimization, pytorch, quantization, tflite, training, transformers
- Archived: false
Getting Started
Install the base package:
python -m pip install optimum
Accelerator integrations may need additional dependencies. See the README and documentation for backend-specific setup and workflows.
Alternatives
- litgpt: LitGPT focuses on configurable LLM training, fine-tuning, evaluation, and serving, while Optimum focuses on hardware-specific optimization across model libraries.
Considerations
- The base install does not by itself provide every accelerator integration. Select the relevant extra and follow its hardware and dependency requirements.
- ONNX integration has moved to the separate optimum-onnx repository, so use that project's installation guidance for ONNX workflows.
- Some paths depend on specific hardware or external runtimes, and their setup differs by provider.
- The repository has 302 open issues, so check current documentation and issue status when assessing a particular integration.
Source repository
Open the original repository on GitHub.
18 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design
October 3, 2026
The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner
October 2, 2026
OOMWOO is an ambitious open-source project enabling users to build their own robot vacuum cleaner using Raspberry Pi, 3D printing, and ROS2. It emphasizes local operation, hackability, and integration with Home Assistant, providing a high-quality, customizable home appliance. This project aims to deliver a fully open hardware, software, and firmware solution for autonomous home cleaning.

Shepherd: Reversible Execution Traces for Programmable Meta-Agents
October 2, 2026
Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents
October 1, 2026
Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.