# VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models

This repository profile is provided by osrepos.com, an open source repository discovery platform.

Source: osrepos.com
Repository profile: https://osrepos.com/repo/verl-project-verl-omni
Generated for open source discovery and AI-assisted research.

VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.

GitHub: https://github.com/verl-project/verl-omni
OSRepos URL: https://osrepos.com/repo/verl-project-verl-omni

## Summary

VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.

## Topics

- diffusion-models
- flow-matching
- multimodal
- reinforcement-learning
- rlhf
- vllm
- Python
- Generative AI

## Repository Information

Last analyzed by OSRepos: Thu Oct 01 2026 09:30:15 GMT+0100 (Western European Summer Time)
Detail views: 0
GitHub clicks: 0

## Safety Notice

OSRepos shares public repositories for knowledge and discovery only. Review source code, dependencies, licenses, and security implications before running or installing anything.

## Content

## Introduction

VeRL-Omni is an advanced reinforcement learning (RL) training framework dedicated to multimodal generative models. It extends the capabilities of the `verl` project, providing a focused environment for the rapid evolution of RL techniques applied to diffusion and omni-modality models. This framework is engineered to simplify, accelerate, and stabilize the training process for cutting-edge generative AI applications.

## Why Use VeRL-Omni and Key Benefits

Multimodal generative RL training presents distinct challenges compared to text-only LLM RL, particularly in model structure, I/O patterns, compute characteristics, and runtime bottlenecks. VeRL-Omni was created to address these specific needs, offering a dedicated platform that can quickly adapt to the evolving landscape of multimodal AI.

VeRL-Omni targets RL post-training for three main families of generative models:

1.  **Diffusion generative models** for image, video, and audio, such as Qwen-Image and Wan2.2.
2.  **Unified multimodal understanding + generation models**, including BAGEL and HunyuanImage-3.0.
3.  **Omni-modality models** that jointly handle text, image, audio, and video, like Qwen3-Omni.

The framework focuses on several key areas to deliver superior performance:

*   **Fast multi-modal rollout**: Utilizes the `vLLM-Omni` backend and optimizes generation through rollout routing, batching, and embed caching.
*   **Flexible & async multi-reward serving**: Supports various multi-reward systems (HPSv3, GenRM-OCR, UnifiedReward) with HTTP scorers and asynchronous reward computation to overlap rollout phases.
*   **Modular training backends**: Offers selectable VeOmni and FSDP2 backends with combinable parallelism (USP/TP/DP) for efficient distributed training.
*   **Stability**: Enhances stability and speed in diffusion RL pipelines via rollout correction and achieves reproducible end-to-end training with deterministic RL.
*   **Efficient and convergent training recipes**: Achieves approximately 25% higher end-to-end throughput on reference setups compared to other implementations, driven by `vLLM-Omni` rollout, FSDP2 trainer, and overlapped reward computation.

## Installation

To get started with VeRL-Omni, please refer to the official [Installation Guide](https://verl-omni.readthedocs.io/en/latest/start/install.html) in the documentation.

## Examples

VeRL-Omni provides robust support for various models and algorithms. An excellent example is optimizing Qwen-Image text rendering accuracy using FlowGRPO, with detailed [recipes](https://github.com/verl-project/verl-omni/tree/main/examples/flowgrpo_trainer/README.md) and [WandB logs](https://wandb.ai/andyzhou/VeRL-Omni-demo/runs/8p8y9olb) available. The project also supports a wide range of models like Qwen-Image, Wan2.2, LTX2.3, MiniMax-H3, Boogu-Image, BAGEL, SD3.5, HunyuanImage-3.0, Qwen3-Omni-Thinker, and Qwen3-TTS, across various modalities and algorithms such as FlowGRPO, Flow-DPPO, DiffusionNFT, DPO, and GSPO.

For specific instructions on running FlowGRPO training on Ascend NPU, consult the [Ascend NPU Quickstart Guide](https://verl-omni.readthedocs.io/en/latest/start/flowgrpo_quickstart_npu.html).

## Links

*   **GitHub Repository**: [verl-project/verl-omni](https://github.com/verl-project/verl-omni){:target="_blank"}
*   **Official Documentation**: [Read the Docs](https://verl-omni.readthedocs.io/en/latest/index.html){:target="_blank"}
*   **DeepWiki**: [Ask DeepWiki.com](https://deepwiki.com/verl-project/verl-omni){:target="_blank"}
*   **Slack Community**: [Join verl Slack](https://join.slack.com/t/verl-project/shared_invite/zt-4bb32of1g-c~szeecZ2fufGeSadsIefw){:target="_blank"}
*   **Project Slides**: [VeRL-Omni Slides](https://drive.google.com/file/d/1EGVFZdvRgSBymMlZa62FObbzEhMoLA0F/view?usp=sharing){:target="_blank"}
*   **WeChat Group**: [Join WeChat Group](docs/assets/WeChat.jpg){:target="_blank"}
*   **Contributing Guide**: [CONTRIBUTING.md](https://github.com/verl-project/verl-omni/blob/main/CONTRIBUTING.md){:target="_blank"}
*   **Citation**: [BibTeX](https://github.com/verl-project/verl-omni#citation-%F0%9F%93%9A){:target="_blank"}