Open Source Deep Learning Projects
Discover 79 open source Deep Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Deep Learning projects here are most often combined with Machine Learning, Python and PyTorch. Last updated October 3, 2026.
79 repositories · updated October 3, 2026

StreamDiffusion: Generate Images in Real Time with Diffusion
StreamDiffusion adapts diffusion pipelines for interactive image generation, with support for text-to-image and image-to-image workflows. It targets developers building responsive GPU-powered demos and applications.

picotron: Train Llama-Like Models with Distributed Parallelism
Picotron is a compact Python framework for learning and experimenting with distributed pre-training of Llama-like models. It demonstrates data, tensor, pipeline, and context parallelism, prioritizing readable code over peak performance.

SyncTalk: Generate Synchronized Talking-Head Videos
SyncTalk is a CVPR 2024 system for generating talking-head videos from a person’s footage and audio. It targets synchronized lip movement, facial expression, and head pose, with workflows for training on a subject or running inference with provided models.

vggt: Reconstruct 3D Scenes from Images
VGGT is a feed-forward vision model that estimates cameras, depth, point maps, and point tracks from one or more images. It suits researchers and developers who need fast 3D scene reconstruction without a traditional multi-stage pipeline.

MuseTalk: Generate Audio-Synced Talking-Head Videos
MuseTalk creates lip-synced video from a source video or image and an audio clip using latent-space inpainting. It supports training and inference workflows, with real-time performance reported on a Tesla V100.

LAM: Create Animatable 3D Gaussian Avatars from One Image
LAM reconstructs a 3D Gaussian head avatar from a single image and supports animation and rendering across devices. It is aimed at developers building digital humans, especially interactive avatars, and requires model assets and a compatible compute environment for local use.

multiresolution-time-series-transformer: Forecast Time Series at Multiple Scales
A PyTorch forecasting model that processes time series at several temporal resolutions, then fuses the representations to predict future values. It suits experimentation with multi-scale forecasting, with adaptations from the paper that users should account for.

PPS-Ctrl: Translate Colonoscopy Images for Depth Estimation
PPS-Ctrl explores controllable sim-to-real translation for colonoscopy images, using per-pixel shading maps to guide Stable Diffusion and ControlNet. The repository currently provides partial pseudocode rather than a complete, ready-to-run implementation.

GPT-SoVITS: Clone Voices and Generate Speech from Text
GPT-SoVITS is a Python toolkit for voice cloning, speech conversion, and text-to-speech. It supports zero-shot synthesis from a short reference clip and fine-tuning with about one minute of voice data, with a WebUI for preparing data and training models.

txtinstruct: Build Instruction-Tuned Models from Your Data
txtinstruct is a Python framework for creating instruction-following datasets and training instruction-tuned models. It is intended for people who want greater control over dataset licensing or to incorporate their own data.

deepface: Analyze and Recognize Faces in Python
DeepFace is a Python library that combines face recognition models and detection tools behind a simple API. Use it to verify identities, search face collections, generate embeddings, or estimate facial attributes.

audio2photoreal: Generate Photorealistic Avatars from Audio
A PyTorch research project for generating photorealistic human face and body motion from conversational audio. It includes person-specific pretrained models, training code, and access to the annotated dataset.