Open Source Deep Learning Projects
Discover 79 open source Deep Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Deep Learning projects here are most often combined with Machine Learning, Python and PyTorch. Last updated October 3, 2026.
79 repositories · updated October 3, 2026

mlx-examples: Explore Machine Learning Models with MLX
A collection of standalone Python examples for the MLX machine learning framework, spanning language, image, video, audio, and multimodal models. Use it to learn MLX or adapt example implementations for experiments on supported hardware.

open-r1: Reproduce DeepSeek-R1 Training and Evaluation
Open R1 is Hugging Face’s toolkit and research project for reproducing the DeepSeek-R1 pipeline with open datasets, training scripts, and evaluation workflows. It is no longer maintained; its training work has moved to TRL.

InfiniteTalk: Generate Audio-Driven Talking Videos
InfiniteTalk generates talking videos from audio and an image or existing video, synchronizing speech with facial expressions and body movement. It is aimed at creators and developers who need dubbed or audio-driven video, including long-form generation.

LlamaFactory: Fine-Tune Large Language and Vision Models
LlamaFactory provides CLI and web interfaces for fine-tuning a broad range of language and vision models. It supports parameter-efficient methods and preference training, with workflows for training, inference, and model export.

pytorch-deep-learning: Learn PyTorch Through Hands-On Lessons
A free, code-first PyTorch course for learning deep learning through notebooks, exercises, and projects. It guides beginners from tensor fundamentals to transfer learning, experiment tracking, and model deployment.

FlashVideo: Generate and Upscale High-Resolution Videos
FlashVideo is a two-stage text-to-video project that generates 270p clips and enhances them to 1080p. It is aimed at researchers and developers exploring efficient high-resolution video generation, with inference code and model weights available.

HunyuanWorld-1.0: Generate Interactive 3D Worlds from Text or Images
HunyuanWorld-1.0 turns text prompts or images into immersive, explorable 3D scenes. It combines panoramic image generation with semantic layering and 3D reconstruction, targeting creators and researchers who need editable world assets.

lance: Store and Query Multimodal Lakehouse Data
Lance is a Rust-based lakehouse format and SDK for AI and machine-learning data. It combines columnar storage with random access, vector and full-text search, and dataset versioning for teams working with multimodal data.

gradio: Build and Share Python Web Apps for Machine Learning
Gradio turns Python functions and machine-learning models into interactive web apps without requiring frontend development. Use it to prototype interfaces, share demos, or build more customized apps with components and event-driven layouts.

Step-Video-T2V: Generate Videos from Text Prompts
Step-Video-T2V is a 30-billion-parameter text-to-video model that generates clips up to 204 frames from English or Chinese prompts. It offers downloadable weights and inference code, but practical use requires substantial GPU memory and multi-GPU setup.

LitServe: Build Custom AI Inference Servers in Python
LitServe is a Python framework for building custom AI inference APIs, from single models to agents and multi-model pipelines. Use it when you need control over request logic, batching, streaming, and deployment rather than a fixed serving abstraction.

Wan2.2: Generate Videos from Text, Images, and Audio
Wan2.2 is a family of open video-generation models for text, image, and audio-driven workflows, plus character animation and replacement. It suits researchers and creators who can run large models on GPU hardware and want local inference options.