Open Source Machine Learning Projects
Discover 191 open source Machine Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Machine Learning projects here are most often combined with Python, Deep Learning and LLM. Last updated October 4, 2026.
191 repositories · updated October 4, 2026

Isaac-GR00T: Adapt Vision-Language Models for Robot Control
NVIDIA Isaac GR00T is a vision-language-action model and toolkit for training and deploying generalist robot skills. It supports inference and fine-tuning across humanoid and other robot embodiments, with GPU hardware and model access requirements.
HunyuanVideo-Avatar: Create Audio-Driven Character Videos
HunyuanVideo-Avatar generates dynamic, emotion-controllable videos of one or more characters from avatar images and audio. It is aimed at creators and researchers who need expressive talking-avatar or dialogue video generation and have access to compatible NVIDIA GPU hardware.

OmniParser: Turn Screenshots into GUI Elements
OmniParser parses interface screenshots into structured elements, helping vision-based agents identify and ground actions on screen. It suits developers building computer-use agents that need visual UI understanding.

clarity-upscaler: Enhance and Upscale Images with AI
Clarity-Upscaler is a Python image-to-image project for increasing image resolution and enhancing details with Stable Diffusion workflows. It suits users comfortable with Cog or image-generation tools who want a configurable alternative to hosted upscaling services.

TextMachina: Build Datasets for Machine-Generated Text Tasks
TextMachina is a Python framework for generating and exploring datasets for machine-generated text detection, attribution, and boundary tasks. It helps researchers and developers combine text sources, language models, and configurable generation pipelines while checking for dataset quality and bias.

CineScale: Generate High-Resolution Video Without Fine-Tuning
CineScale is an inference framework for generating high-resolution video with pretrained diffusion models, without fine-tuning. It targets researchers and practitioners who want to upscale generation beyond a model’s training resolution, including 4K workflows.

scikit-learn: Build Machine Learning Models in Python
scikit-learn is a Python machine learning library built on SciPy. It helps developers and data practitioners build and evaluate machine learning workflows, with plotting support and documentation for installation and development.

ggml: Build Portable Tensor Workloads for Machine Learning
ggml is a dependency-free C/C++ tensor library for building machine-learning workloads across CPUs and other backends. It is suited to developers who need low-level control over tensor operations, quantization, and deployment across different platforms.

pedalboard: Process Audio and Apply Effects in Python
Pedalboard is a Python library for reading, transforming, and rendering audio, with built-in effects and support for VST3 and Audio Unit plugins. It suits audio developers and machine-learning workflows that need programmable audio processing.

StreamDiffusion: Generate Images in Real Time with Diffusion
StreamDiffusion adapts diffusion pipelines for interactive image generation, with support for text-to-image and image-to-image workflows. It targets developers building responsive GPU-powered demos and applications.

Toolkit-for-Prompt-Compression: Evaluate and Apply Prompt Compression
PCToolkit is a Python toolkit for applying and evaluating prompt-compression methods for large language models. It brings five compressors, datasets, and evaluation metrics behind modular interfaces, making it useful for comparing methods across language tasks.

picotron: Train Llama-Like Models with Distributed Parallelism
Picotron is a compact Python framework for learning and experimenting with distributed pre-training of Llama-like models. It demonstrates data, tensor, pipeline, and context parallelism, prioritizing readable code over peak performance.