Open Source Diffusion Models
Diffusion models are generative machine-learning systems that learn to create or transform data by gradually removing noise. Guided by inputs such as text, images, or audio, they can produce images and video, edit existing media, and support tasks such as pose or appearance changes. These methods offer flexible ways to generate complex visual and other media, while allowing control over the output through prompts, conditioning, and model settings.
Open source tools in this area include pretrained model weights, inference libraries, training and fine-tuning frameworks, and interfaces for generating or editing media. When choosing, check the license and permitted uses, model quality for your task, hardware and memory requirements, maintenance activity, and compatibility with your existing workflow. These tools are useful to developers, researchers, artists, and organizations building or studying generative media systems.
8 repositories · updated October 1, 2026

VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models
VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.

Transformer Lab App: An Open Source Platform for Frontier AI/ML Workflows
Transformer Lab App is an open-source machine learning research platform designed for frontier AI/ML workflows. It provides a comprehensive toolkit for large language models, allowing users to train, tune, and chat on their own machines, whether locally, on-prem, or in the cloud. Backed by Mozilla, this cross-platform application simplifies experimentation with a wide range of models.
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation
HunyuanVideo-Avatar is a cutting-edge project by Tencent-Hunyuan for high-fidelity, audio-driven human animation. Utilizing a multimodal diffusion transformer, it generates dynamic, emotion-controllable, and multi-character dialogue videos. This innovative system addresses critical challenges in character consistency, emotion alignment, and multi-character animation, making it suitable for diverse applications like e-commerce and social media.

CineScale: Unlocking 4K High-Resolution Cinematic Video Generation
CineScale is an innovative GitHub repository by Eyeline-Labs, extending FreeScale to enable high-resolution cinematic video generation. It provides models and tools to achieve up to 4K video output, leveraging diffusion models for advanced visual content creation. This project offers a robust framework for researchers and developers to generate stunning, high-definition videos.

StreamDiffusion: Real-Time Interactive Generation with Diffusion Pipelines
StreamDiffusion is an innovative diffusion pipeline designed for real-time interactive generation, significantly enhancing the performance of current diffusion-based image generation techniques. It offers a pipeline-level solution to achieve high-speed image and text-to-image generation, making interactive AI experiences more accessible. This project introduces several key features to optimize computational efficiency and GPU utilization.

FlashVideo: Efficient High-Resolution Video Generation with Flowing Fidelity
FlashVideo is an innovative GitHub repository that introduces a novel approach for efficient high-resolution video generation. It leverages a two-stage diffusion model to produce detailed videos, scaling from 270p to 1080p. This project focuses on maintaining fidelity to detail while significantly improving the efficiency of the video generation process.

Step-Video-T2V: State-of-the-Art Text-to-Video Generation Model
Step-Video-T2V is a state-of-the-art text-to-video pre-trained model capable of generating videos up to 204 frames with 30 billion parameters. It achieves high efficiency through a deep compression Video-VAE and enhances visual quality using Direct Preference Optimization (DPO). The model's performance is validated on its novel benchmark, Step-Video-T2V-Eval, demonstrating superior text-to-video quality.

Leffa: Controllable Person Image Generation with Flow Fields in Attention
Leffa is a unified framework for controllable person image generation, enabling precise manipulation of appearance through virtual try-on and pose via pose transfer. This project addresses the common issue of fine-grained textural detail distortion by learning flow fields in attention, guiding target queries to correct reference keys. It achieves state-of-the-art performance, maintaining high image quality while significantly reducing detail distortion.