{"name":"VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models","description":"VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.","github":"https://github.com/verl-project/verl-omni","url":"https://osrepos.com/repo/verl-project-verl-omni","source":"osrepos.com","sourceDescription":"This repository profile is provided by osrepos.com, an open source repository discovery platform.","repositoryProfile":"https://osrepos.com/repo/verl-project-verl-omni","generatedFor":"open source discovery and AI-assisted research","markdown":"https://osrepos.com/repo/verl-project-verl-omni.md","json":"https://osrepos.com/repo/verl-project-verl-omni.json","topics":["diffusion-models","flow-matching","multimodal","reinforcement-learning","rlhf","vllm","Python","Generative AI"],"keywords":["diffusion-models","flow-matching","multimodal","reinforcement-learning","rlhf","vllm","Python","Generative AI"],"stars":null,"summary":"VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.","content":"## Introduction\n\nVeRL-Omni is an advanced reinforcement learning (RL) training framework dedicated to multimodal generative models. It extends the capabilities of the `verl` project, providing a focused environment for the rapid evolution of RL techniques applied to diffusion and omni-modality models. This framework is engineered to simplify, accelerate, and stabilize the training process for cutting-edge generative AI applications.\n\n## Why Use VeRL-Omni and Key Benefits\n\nMultimodal generative RL training presents distinct challenges compared to text-only LLM RL, particularly in model structure, I/O patterns, compute characteristics, and runtime bottlenecks. VeRL-Omni was created to address these specific needs, offering a dedicated platform that can quickly adapt to the evolving landscape of multimodal AI.\n\nVeRL-Omni targets RL post-training for three main families of generative models:\n\n1.  **Diffusion generative models** for image, video, and audio, such as Qwen-Image and Wan2.2.\n2.  **Unified multimodal understanding + generation models**, including BAGEL and HunyuanImage-3.0.\n3.  **Omni-modality models** that jointly handle text, image, audio, and video, like Qwen3-Omni.\n\nThe framework focuses on several key areas to deliver superior performance:\n\n*   **Fast multi-modal rollout**: Utilizes the `vLLM-Omni` backend and optimizes generation through rollout routing, batching, and embed caching.\n*   **Flexible & async multi-reward serving**: Supports various multi-reward systems (HPSv3, GenRM-OCR, UnifiedReward) with HTTP scorers and asynchronous reward computation to overlap rollout phases.\n*   **Modular training backends**: Offers selectable VeOmni and FSDP2 backends with combinable parallelism (USP/TP/DP) for efficient distributed training.\n*   **Stability**: Enhances stability and speed in diffusion RL pipelines via rollout correction and achieves reproducible end-to-end training with deterministic RL.\n*   **Efficient and convergent training recipes**: Achieves approximately 25% higher end-to-end throughput on reference setups compared to other implementations, driven by `vLLM-Omni` rollout, FSDP2 trainer, and overlapped reward computation.\n\n## Installation\n\nTo get started with VeRL-Omni, please refer to the official [Installation Guide](https://verl-omni.readthedocs.io/en/latest/start/install.html) in the documentation.\n\n## Examples\n\nVeRL-Omni provides robust support for various models and algorithms. An excellent example is optimizing Qwen-Image text rendering accuracy using FlowGRPO, with detailed [recipes](https://github.com/verl-project/verl-omni/tree/main/examples/flowgrpo_trainer/README.md) and [WandB logs](https://wandb.ai/andyzhou/VeRL-Omni-demo/runs/8p8y9olb) available. The project also supports a wide range of models like Qwen-Image, Wan2.2, LTX2.3, MiniMax-H3, Boogu-Image, BAGEL, SD3.5, HunyuanImage-3.0, Qwen3-Omni-Thinker, and Qwen3-TTS, across various modalities and algorithms such as FlowGRPO, Flow-DPPO, DiffusionNFT, DPO, and GSPO.\n\nFor specific instructions on running FlowGRPO training on Ascend NPU, consult the [Ascend NPU Quickstart Guide](https://verl-omni.readthedocs.io/en/latest/start/flowgrpo_quickstart_npu.html).\n\n## Links\n\n*   **GitHub Repository**: [verl-project/verl-omni](https://github.com/verl-project/verl-omni){:target=\"_blank\"}\n*   **Official Documentation**: [Read the Docs](https://verl-omni.readthedocs.io/en/latest/index.html){:target=\"_blank\"}\n*   **DeepWiki**: [Ask DeepWiki.com](https://deepwiki.com/verl-project/verl-omni){:target=\"_blank\"}\n*   **Slack Community**: [Join verl Slack](https://join.slack.com/t/verl-project/shared_invite/zt-4bb32of1g-c~szeecZ2fufGeSadsIefw){:target=\"_blank\"}\n*   **Project Slides**: [VeRL-Omni Slides](https://drive.google.com/file/d/1EGVFZdvRgSBymMlZa62FObbzEhMoLA0F/view?usp=sharing){:target=\"_blank\"}\n*   **WeChat Group**: [Join WeChat Group](docs/assets/WeChat.jpg){:target=\"_blank\"}\n*   **Contributing Guide**: [CONTRIBUTING.md](https://github.com/verl-project/verl-omni/blob/main/CONTRIBUTING.md){:target=\"_blank\"}\n*   **Citation**: [BibTeX](https://github.com/verl-project/verl-omni#citation-%F0%9F%93%9A){:target=\"_blank\"}","metrics":{"detailViews":0,"githubClicks":0},"dates":{"published":null,"modified":"2026-10-01T08:30:15.000Z"}}