optimum: Optimize Model Training and Inference on Target Hardware

Summary
Hugging Face Optimum adds tools for optimizing model training and inference across hardware backends. It suits teams using Transformers, Diffusers, TIMM, or Sentence Transformers who need hardware-specific deployment or training workflows.
At a glance
- Language
- Python
- License
- Apache-2.0
- Stars
- 3.5k
- Forks
- 695
- Added to OSRepos
- January 6, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Optimum is a Python library that connects Hugging Face model workflows with optimization and acceleration tools for different hardware. It helps developers export, run, and train models using backends such as OpenVINO, ONNX Runtime, ExecuTorch, and accelerator-specific integrations.
It is most useful when a standard Transformers workflow needs to target a particular device or inference runtime. The core package is a starting point, while many backends require separate extras or companion projects.
Key Features
- Integrates optimization workflows with Transformers, Diffusers, TIMM, and Sentence Transformers.
- Supports model export and optimized inference across several runtimes and hardware ecosystems.
- Provides integrations for Intel OpenVINO, Intel Gaudi, AWS Inferentia and Trainium, AMD hardware, NVIDIA TensorRT-LLM, and other providers.
- Offers training integrations for supported accelerators, including Gaudi and AWS Trainium.
- Includes command-line and programmatic workflows for supported export and optimization tasks.
- Connects users to separate projects for some capabilities, including Optimum ONNX, Optimum Quanto, and Optimum ExecuTorch.
Use Cases
- A machine-learning engineer can export a Transformers model for an inference runtime supported by their deployment hardware.
- An application team can evaluate OpenVINO or another supported backend when deploying models on a specific accelerator ecosystem.
- A model trainer can use the Gaudi or Trainium integrations when training on supported hardware.
- An edge developer can use the Optimum ExecuTorch project to export Transformers models for PyTorch-based on-device inference.
Project Facts
- Language: Python
- License: Apache-2.0
- Stars: 3.5k
- Forks: 695
- Topics: graphcore, habana, inference, intel, onnx, onnxruntime, optimization, pytorch, quantization, tflite, training, transformers
- Archived: false
Getting Started
Install the base package:
python -m pip install optimum
Accelerator integrations may need additional dependencies. See the README and documentation for backend-specific setup and workflows.
Alternatives
- litgpt: LitGPT focuses on configurable LLM training, fine-tuning, evaluation, and serving, while Optimum focuses on hardware-specific optimization across model libraries.
Considerations
- The base install does not by itself provide every accelerator integration. Select the relevant extra and follow its hardware and dependency requirements.
- ONNX integration has moved to the separate optimum-onnx repository, so use that project's installation guidance for ONNX workflows.
- Some paths depend on specific hardware or external runtimes, and their setup differs by provider.
- The repository has 302 open issues, so check current documentation and issue status when assessing a particular integration.
Source repository
Open the original repository on GitHub.
18 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.

agent-anvil: Test AI Agent Tool Use in CI
October 1, 2026
Agent Anvil evaluates tool-using AI agents through scenario-based runs, trace checks, and optional semantic grading. It helps teams catch unsafe or incorrect tool behavior before it reaches production.