Open Source Generative AI Projects
Discover 80 open source Generative AI repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Generative AI projects here are most often combined with Python, Machine Learning and AI. Last updated October 4, 2026.
80 repositories · updated October 4, 2026

riffusion-hobby: Generate Music Audio with Stable Diffusion
Riffusion-hobby is a Python library and application for generating music and audio with stable diffusion, including prompt interpolation and spectrogram-to-audio conversion. It suits developers and experimenters who want to run inference locally, but is no longer actively maintained.

audio2photoreal: Generate Photorealistic Avatars from Audio
A PyTorch research project for generating photorealistic human face and body motion from conversational audio. It includes person-specific pretrained models, training code, and access to the annotated dataset.

mlx-examples: Explore Machine Learning Models with MLX
A collection of standalone Python examples for the MLX machine learning framework, spanning language, image, video, audio, and multimodal models. Use it to learn MLX or adapt example implementations for experiments on supported hardware.

open-r1: Reproduce DeepSeek-R1 Training and Evaluation
Open R1 is Hugging Face’s toolkit and research project for reproducing the DeepSeek-R1 pipeline with open datasets, training scripts, and evaluation workflows. It is no longer maintained; its training work has moved to TRL.

InfiniteTalk: Generate Audio-Driven Talking Videos
InfiniteTalk generates talking videos from audio and an image or existing video, synchronizing speech with facial expressions and body movement. It is aimed at creators and developers who need dubbed or audio-driven video, including long-form generation.

llama-cpp-python: Run llama.cpp Models from Python
Python bindings for llama.cpp let developers run GGUF language models locally through a high-level API, low-level C bindings, or an OpenAI-compatible server. Useful when you want local inference and control over hardware backends.

podcastfy: Turn Multimodal Sources Into AI Podcasts
Podcastfy is a Python package and CLI for turning websites, PDFs, images, YouTube videos, and topics into multilingual conversational audio. It suits developers and creators who want to customize or automate podcast generation with hosted or local language models.

FlashVideo: Generate and Upscale High-Resolution Videos
FlashVideo is a two-stage text-to-video project that generates 270p clips and enhances them to 1080p. It is aimed at researchers and developers exploring efficient high-resolution video generation, with inference code and model weights available.

HunyuanWorld-1.0: Generate Interactive 3D Worlds from Text or Images
HunyuanWorld-1.0 turns text prompts or images into immersive, explorable 3D scenes. It combines panoramic image generation with semantic layering and 3D reconstruction, targeting creators and researchers who need editable world assets.

factorio-blueprint-visualizer: Turn Blueprints into Artwork
A browser-based tool and JavaScript library that turns Factorio blueprints into customizable vector illustrations. Use it to explore factory layouts, create images for sharing, or prepare designs for pen plotting.

Step-Video-T2V: Generate Videos from Text Prompts
Step-Video-T2V is a 30-billion-parameter text-to-video model that generates clips up to 204 frames from English or Chinese prompts. It offers downloadable weights and inference code, but practical use requires substantial GPU memory and multi-GPU setup.

Wan2.2: Generate Videos from Text, Images, and Audio
Wan2.2 is a family of open video-generation models for text, image, and audio-driven workflows, plus character animation and replacement. It suits researchers and creators who can run large models on GPU hardware and want local inference options.