Open Source Machine Learning Projects
Discover 191 open source Machine Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Machine Learning projects here are most often combined with Python, Deep Learning and LLM. Last updated October 4, 2026.
191 repositories · updated October 4, 2026

maestro: Fine-Tune Vision-Language Models
Maestro streamlines fine-tuning for multimodal vision-language models with ready-to-use recipes, a CLI, and a Python API. It suits developers adapting supported models to tasks such as JSON extraction and object detection.

KBLaM: Add Knowledge Bases to Language Models
KBLaM is a research implementation for giving transformer language models access to external knowledge through learned adapters and special knowledge tokens. It is aimed at researchers testing knowledge-grounded answers without a separate retrieval module.

MonoPCC: Estimate Monocular Depth in Endoscopic Images
MonoPCC is a PyTorch method for self-supervised monocular depth estimation on endoscopic images, using a photometric-invariant cycle constraint. It is aimed at researchers reproducing or extending depth and pose experiments on surgical-video datasets.

spotlight: Build PyTorch Recommender Models
Spotlight is a Python library for building and testing recommender models with PyTorch. It includes factorization and sequence-model building blocks, ranking losses, and dataset utilities for research and prototyping.

fast-music-remover: Remove Music and Noise from Media Audio
Fast Music Remover processes audio from uploaded files or internet media to reduce background music and noise. It combines a C++ audio processor using DeepFilterNet with a Flask web interface, and can be run in Docker or set up manually.

PETSA: Adapt Time-Series Forecasters at Test Time
PETSA is a parameter-efficient test-time adaptation method for time-series forecasting. It updates small input and output calibration modules rather than the full model, targeting non-stationary data with lower adaptation costs.

insanely-fast-whisper: Transcribe Audio with Whisper
A command-line tool for running Whisper speech recognition on your own NVIDIA GPU or Apple Silicon Mac. It combines batching and optional Flash Attention to speed up transcription, with options for timestamps, translation, and speaker diarization.

flash-attention: Compute Exact Attention Faster and with Less Memory
FlashAttention provides GPU-optimized implementations of exact attention for deep learning, reducing memory use while improving speed. It is aimed at practitioners training or serving Transformer models who can build and run GPU kernels.

verifiers: Build Environments for LLM Training and Evaluation
verifiers is a Python library for building environments to train and evaluate large language models. It fits teams developing LLM reinforcement-learning workflows or reusable evaluations, especially those using Prime Intellect's training tools.

gaussian-splatting: Reconstruct and Render 3D Scenes in Real Time
The authors’ reference implementation of 3D Gaussian Splatting, which reconstructs scenes from posed images and renders novel viewpoints. It suits graphics and vision researchers and practitioners with CUDA-capable GPUs who need a trainable, interactive scene representation.

argo-workflows: Run Container Workflows on Kubernetes
Argo Workflows runs container-based jobs on Kubernetes as step sequences or dependency graphs. It suits teams building batch, data, machine-learning, and CI/CD pipelines that need Kubernetes-native execution and orchestration.

LLMSanitize: Detect Contamination in NLP Data and LLMs
LLMSanitize brings together methods for checking whether NLP datasets or language models may be contaminated by training data. It is aimed at researchers and evaluators who need to assess benchmark reliability across open- and closed-data settings.