Open Source Deep Learning Projects
Discover 73 open source Deep Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Deep Learning projects here are most often combined with Python, Machine Learning and AI. Last updated October 3, 2026.
73 repositories · updated October 3, 2026

VoxCPM: Generate and Clone Multilingual Speech
VoxCPM is a tokenizer-free text-to-speech system for multilingual speech generation, voice design, and voice cloning. Its VoxCPM2 release targets teams and developers who need expressive speech synthesis and can support a 2B-parameter model.

ai-engineering-from-scratch: Learn AI by Building It
A free, MIT-licensed curriculum that teaches AI engineering through runnable lessons and hands-on artifacts. It covers foundations through LLMs, agents, protocols, and production, with paths for beginners and experienced developers.

JARVIS: Connect Language Models to AI Tools
JARVIS, also known as HuggingGPT, uses an LLM to plan requests, choose expert models from Hugging Face, run tasks, and combine their results. It suits research and prototyping that explores natural-language orchestration of AI models.

peft: Fine-Tune Models with Fewer Trainable Parameters
Hugging Face PEFT adapts pretrained models by training a small set of additional parameters instead of updating the full model. It is for practitioners who want to reduce fine-tuning compute and storage costs across supported model workflows.

Qwen3: Alibaba Cloud's Advanced Large Language Model Series
Qwen3 is a powerful series of large language models developed by the Qwen team at Alibaba Cloud. It offers advanced capabilities in reasoning, multilingual support, and long-context understanding, available in various sizes and modes for diverse applications. This repository provides comprehensive resources for running, deploying, and building with Qwen3 models.

PyTorch Image Models (timm): The Ultimate Collection of Image Encoders
PyTorch Image Models (timm) is an extensive library offering the largest collection of PyTorch image encoders and backbones. It provides a wide array of state-of-the-art models, complete with pretrained weights, training, evaluation, and inference scripts. This makes it an invaluable resource for researchers and developers working with computer vision tasks in PyTorch.

kapre: Add Audio Preprocessing Layers to Keras Models
Kapre provides TensorFlow-backed Keras layers for transforming audio waveforms into STFT, mel-spectrogram, and related features inside a model. It suits audio ML developers who want preprocessing to travel with training and deployment.

Kimi-k1.5: Scaling Reinforcement Learning with LLMs and Multimodality
Kimi-k1.5 introduces an o1-level multi-modal model that significantly advances reinforcement learning with Large Language Models. It demonstrates state-of-the-art performance in short-CoT tasks, outperforming leading models like GPT-4o and Claude Sonnet 3.5, and matches o1 performance in long-CoT scenarios across various modalities. This project highlights key innovations in long context scaling and improved policy optimization.
CoTracker: A Powerful Model for Tracking Any Point on a Video
CoTracker is a state-of-the-art model developed by Facebook AI Research and the University of Oxford, designed for tracking any point (pixel) across video sequences. This transformer-based solution offers fast, accurate, and quasi-dense point tracking capabilities. It is an invaluable tool for researchers and developers in computer vision, enabling precise analysis of motion in videos.

notebooks: Learn and Apply Computer Vision Models
Roboflow notebooks is a hands-on tutorial collection for computer vision, covering model training, inference, detection, segmentation, and related tasks. Use it to explore techniques and run examples in hosted notebook environments.

Spark-TTS: Generate Speech and Clone Voices from Text
Spark-TTS is a PyTorch inference project for bilingual text-to-speech and zero-shot voice cloning. It uses a Qwen2.5-based model to generate speech and supports adjustable voice characteristics.

TRELLIS: Generate 3D Assets from Text or Images
TRELLIS is a research model and toolkit for generating 3D assets from text or images. Its structured latent representation can produce meshes, 3D Gaussians, and radiance fields, but local use requires a compatible NVIDIA GPU and a substantial setup.