VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models
This repository profile is provided by osrepos.com, an open source repository discovery platform.

Summary
VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.
Repository Information
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Introduction
VeRL-Omni is an advanced reinforcement learning (RL) training framework dedicated to multimodal generative models. It extends the capabilities of the verl project, providing a focused environment for the rapid evolution of RL techniques applied to diffusion and omni-modality models. This framework is engineered to simplify, accelerate, and stabilize the training process for cutting-edge generative AI applications.
Why Use VeRL-Omni and Key Benefits
Multimodal generative RL training presents distinct challenges compared to text-only LLM RL, particularly in model structure, I/O patterns, compute characteristics, and runtime bottlenecks. VeRL-Omni was created to address these specific needs, offering a dedicated platform that can quickly adapt to the evolving landscape of multimodal AI.
VeRL-Omni targets RL post-training for three main families of generative models:
- Diffusion generative models for image, video, and audio, such as Qwen-Image and Wan2.2.
- Unified multimodal understanding + generation models, including BAGEL and HunyuanImage-3.0.
- Omni-modality models that jointly handle text, image, audio, and video, like Qwen3-Omni.
The framework focuses on several key areas to deliver superior performance:
- Fast multi-modal rollout: Utilizes the
vLLM-Omnibackend and optimizes generation through rollout routing, batching, and embed caching. - Flexible & async multi-reward serving: Supports various multi-reward systems (HPSv3, GenRM-OCR, UnifiedReward) with HTTP scorers and asynchronous reward computation to overlap rollout phases.
- Modular training backends: Offers selectable VeOmni and FSDP2 backends with combinable parallelism (USP/TP/DP) for efficient distributed training.
- Stability: Enhances stability and speed in diffusion RL pipelines via rollout correction and achieves reproducible end-to-end training with deterministic RL.
- Efficient and convergent training recipes: Achieves approximately 25% higher end-to-end throughput on reference setups compared to other implementations, driven by
vLLM-Omnirollout, FSDP2 trainer, and overlapped reward computation.
Installation
To get started with VeRL-Omni, please refer to the official Installation Guide in the documentation.
Examples
VeRL-Omni provides robust support for various models and algorithms. An excellent example is optimizing Qwen-Image text rendering accuracy using FlowGRPO, with detailed recipes and WandB logs available. The project also supports a wide range of models like Qwen-Image, Wan2.2, LTX2.3, MiniMax-H3, Boogu-Image, BAGEL, SD3.5, HunyuanImage-3.0, Qwen3-Omni-Thinker, and Qwen3-TTS, across various modalities and algorithms such as FlowGRPO, Flow-DPPO, DiffusionNFT, DPO, and GSPO.
For specific instructions on running FlowGRPO training on Ascend NPU, consult the Ascend NPU Quickstart Guide.
Links
- GitHub Repository: verl-project/verl-omni
- Official Documentation: Read the Docs
- DeepWiki: Ask DeepWiki.com
- Slack Community: Join verl Slack
- Project Slides: VeRL-Omni Slides
- WeChat Group: Join WeChat Group
- Contributing Guide: CONTRIBUTING.md
- Citation: BibTeX
Related repositories
Similar repositories that may be relevant next.

CineScale: Unlocking 4K High-Resolution Cinematic Video Generation
December 18, 2025
CineScale is an innovative GitHub repository by Eyeline-Labs, extending FreeScale to enable high-resolution cinematic video generation. It provides models and tools to achieve up to 4K video output, leveraging diffusion models for advanced visual content creation. This project offers a robust framework for researchers and developers to generate stunning, high-definition videos.

FlashVideo: Efficient High-Resolution Video Generation with Flowing Fidelity
November 5, 2025
FlashVideo is an innovative GitHub repository that introduces a novel approach for efficient high-resolution video generation. It leverages a two-stage diffusion model to produce detailed videos, scaling from 270p to 1080p. This project focuses on maintaining fidelity to detail while significantly improving the efficiency of the video generation process.
Source repository
Open the original repository on GitHub.