Open Source PyTorch Projects
Discover 46 open source PyTorch repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. PyTorch projects here are most often combined with Machine Learning, Python and Deep Learning. Last updated October 3, 2026.
46 repositories · updated October 3, 2026

audio2photoreal: Generate Photorealistic Avatars from Audio
A PyTorch research project for generating photorealistic human face and body motion from conversational audio. It includes person-specific pretrained models, training code, and access to the annotated dataset.

InfiniteTalk: Generate Audio-Driven Talking Videos
InfiniteTalk generates talking videos from audio and an image or existing video, synchronizing speech with facial expressions and body movement. It is aimed at creators and developers who need dubbed or audio-driven video, including long-form generation.

pytorch-deep-learning: Learn PyTorch Through Hands-On Lessons
A free, code-first PyTorch course for learning deep learning through notebooks, exercises, and projects. It guides beginners from tensor fundamentals to transfer learning, experiment tracking, and model deployment.

FlashVideo: Generate and Upscale High-Resolution Videos
FlashVideo is a two-stage text-to-video project that generates 270p clips and enhances them to 1080p. It is aimed at researchers and developers exploring efficient high-resolution video generation, with inference code and model weights available.

text-generation-inference: Serve Large Language Models
A toolkit for serving large language models through text-generation APIs. It provides GPU-oriented inference features such as continuous batching, token streaming, tensor parallelism, and quantization, but is now in maintenance mode.

Step-Video-T2V: Generate Videos from Text Prompts
Step-Video-T2V is a 30-billion-parameter text-to-video model that generates clips up to 204 frames from English or Chinese prompts. It offers downloadable weights and inference code, but practical use requires substantial GPU memory and multi-GPU setup.

Wan2.2: Generate Videos from Text, Images, and Audio
Wan2.2 is a family of open video-generation models for text, image, and audio-driven workflows, plus character animation and replacement. It suits researchers and creators who can run large models on GPU hardware and want local inference options.

FinGPT: Adapt Language Models for Financial Tasks
FinGPT provides financial datasets, fine-tuned language models, benchmarks, and workflows for tasks such as sentiment analysis and forecasting. It suits researchers and developers adapting open models to finance, with local GPU inference or supported cloud APIs.

LivePortrait: Animate Portraits from Images and Video
LivePortrait is a PyTorch project for animating human and animal portraits using a driving video or motion template. It supports portrait video editing and offers command-line inference and a Gradio interface.

Leffa: Generate Controllable Person Images
Leffa is a diffusion-based framework for virtual try-on and pose transfer that aims to preserve fine-grained details from reference images. It is suited to researchers and developers building or evaluating controllable person-image generation systems.