VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models

This repository profile is provided by osrepos.com, an open source repository discovery platform.

VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models

Summary

VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.

Repository Information

Analyzed by OSRepos on October 1, 2026

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Introduction

VeRL-Omni is an advanced reinforcement learning (RL) training framework dedicated to multimodal generative models. It extends the capabilities of the verl project, providing a focused environment for the rapid evolution of RL techniques applied to diffusion and omni-modality models. This framework is engineered to simplify, accelerate, and stabilize the training process for cutting-edge generative AI applications.

Why Use VeRL-Omni and Key Benefits

Multimodal generative RL training presents distinct challenges compared to text-only LLM RL, particularly in model structure, I/O patterns, compute characteristics, and runtime bottlenecks. VeRL-Omni was created to address these specific needs, offering a dedicated platform that can quickly adapt to the evolving landscape of multimodal AI.

VeRL-Omni targets RL post-training for three main families of generative models:

  • Diffusion generative models for image, video, and audio, such as Qwen-Image and Wan2.2.
  • Unified multimodal understanding + generation models, including BAGEL and HunyuanImage-3.0.
  • Omni-modality models that jointly handle text, image, audio, and video, like Qwen3-Omni.

The framework focuses on several key areas to deliver superior performance:

  • Fast multi-modal rollout: Utilizes the vLLM-Omni backend and optimizes generation through rollout routing, batching, and embed caching.
  • Flexible & async multi-reward serving: Supports various multi-reward systems (HPSv3, GenRM-OCR, UnifiedReward) with HTTP scorers and asynchronous reward computation to overlap rollout phases.
  • Modular training backends: Offers selectable VeOmni and FSDP2 backends with combinable parallelism (USP/TP/DP) for efficient distributed training.
  • Stability: Enhances stability and speed in diffusion RL pipelines via rollout correction and achieves reproducible end-to-end training with deterministic RL.
  • Efficient and convergent training recipes: Achieves approximately 25% higher end-to-end throughput on reference setups compared to other implementations, driven by vLLM-Omni rollout, FSDP2 trainer, and overlapped reward computation.

Installation

To get started with VeRL-Omni, please refer to the official Installation Guide in the documentation.

Examples

VeRL-Omni provides robust support for various models and algorithms. An excellent example is optimizing Qwen-Image text rendering accuracy using FlowGRPO, with detailed recipes and WandB logs available. The project also supports a wide range of models like Qwen-Image, Wan2.2, LTX2.3, MiniMax-H3, Boogu-Image, BAGEL, SD3.5, HunyuanImage-3.0, Qwen3-Omni-Thinker, and Qwen3-TTS, across various modalities and algorithms such as FlowGRPO, Flow-DPPO, DiffusionNFT, DPO, and GSPO.

For specific instructions on running FlowGRPO training on Ascend NPU, consult the Ascend NPU Quickstart Guide.

Links

Related repositories

Similar repositories that may be relevant next.

Source repository

Open the original repository on GitHub.

View on GitHub
OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️