TensorRT-LLM: Optimize LLM Inference on NVIDIA GPUs

TensorRT-LLM: Optimize LLM Inference on NVIDIA GPUs

Summary

TensorRT-LLM is a Python framework and runtime for efficient LLM and visual-generation inference on NVIDIA GPUs. It suits teams deploying models on NVIDIA hardware that need optimized kernels and configurable single- or distributed-GPU execution.

At a glance

Language
Python
License
NOASSERTION
Stars
14.8k
Forks
2.8k
Added to OSRepos
July 3, 2026
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

TensorRT-LLM is a framework for running large language models and visual-generation models efficiently on NVIDIA GPUs. It combines optimized inference operations with a runtime for orchestrating execution, addressing the performance and deployment needs of GPU-based inference workloads.

It is aimed at developers and infrastructure teams who want to tune or customize inference rather than use a hardware-agnostic serving layer. Its Python-native design supports model and runtime customization, while its APIs cover deployments from one GPU to multiple GPUs or nodes.

Key Features

  • Specialized kernels for common inference operations, including attention, GEMMs, and mixture-of-experts workloads.
  • Runtime optimizations including speculative decoding and prefill-decode disaggregation.
  • Python API for configuring and running supported models.
  • Support for single-GPU, multi-GPU, and multi-node inference setups.
  • Built-in parallelism strategies for distributing inference work.
  • PyTorch-native components that developers can modify and extend.
  • Integration with NVIDIA Dynamo and Triton Inference Server.

Use Cases

  • Inference platform teams serving LLMs on NVIDIA GPUs that need to tune throughput or latency for their workloads.
  • Developers adapting supported model implementations or runtime behavior using Python and PyTorch.
  • Operators deploying large models across multiple GPUs or nodes when a single GPU is insufficient.
  • Teams evaluating optimizations such as speculative decoding or prefill-decode disaggregation for their serving setup.
  • Visual-generation workloads on NVIDIA GPUs that can use the project's supported inference components.

Project Facts

  • Language: Python
  • License: NOASSERTION
  • Stars: 14.8k
  • Forks: 2.8k
  • Topics: blackwell, cuda, llm-serving, moe, pytorch
  • Archived: No

Getting Started

Start with the Quick Start Guide, then check the Installation Guide and supported hardware and models. See the repository README for project details and additional links.

Alternatives

  • Text Generation Inference: Text Generation Inference is a general-purpose LLM serving toolkit, while TensorRT-LLM focuses on NVIDIA GPU-optimized inference and distributed execution.

Considerations

  • The project targets NVIDIA GPUs, so it is not a general-purpose GPU-agnostic inference framework.
  • Multi-GPU and multi-node use introduces deployment and configuration complexity; check the support matrix for hardware and model compatibility.
  • The repository metadata reports the license as NOASSERTION. Review the repository's license information before adopting it.
  • The README says anonymous telemetry is collected by default and documents ways to opt out, including environment variables, a configuration file, a Python option, and CLI flags.

Source repository

Open the original repository on GitHub.

29 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design

web-design: A Claude Code SKILL for Spec-First Web Page Design

October 3, 2026

The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

Claude CodeClaude SkillDesign System
OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner

OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner

October 2, 2026

OOMWOO is an ambitious open-source project enabling users to build their own robot vacuum cleaner using Raspberry Pi, 3D printing, and ROS2. It emphasizes local operation, hackability, and integration with Home Assistant, providing a high-quality, customizable home appliance. This project aims to deliver a fully open hardware, software, and firmware solution for autonomous home cleaning.

RoboticsOpen SourceDiy
Shepherd: Reversible Execution Traces for Programmable Meta-Agents

Shepherd: Reversible Execution Traces for Programmable Meta-Agents

October 2, 2026

Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

PythonAIAgent Framework
Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

October 1, 2026

Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

PythonAIAgent
OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️