Open Source Machine Learning Projects
Discover 191 open source Machine Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Machine Learning projects here are most often combined with Python, Deep Learning and LLM. Last updated October 4, 2026.
191 repositories · updated October 4, 2026

GLM-5: Power Long-Horizon Coding and Agentic Tasks
GLM-5 is a family of large language models from Z.ai for complex coding, systems engineering, and long-running agent tasks. The repository provides model downloads, serving guidance, and fine-tuning references.

bg-remove: Remove Image Backgrounds in Your Browser
bg-remove is a browser-based image background remover that runs machine-learning models locally with Transformers.js. It suits people who want to edit images without uploading them to a server, with optional WebGPU acceleration where supported.

physicsnemo: Build and Train Physics AI Models
NVIDIA PhysicsNeMo is a PyTorch framework for building and training machine-learning models for physics and engineering. It combines reusable model components with end-to-end recipes for scientific data and workloads.

Qwen3-VL: Understand Images, Video, and Text with Multimodal Models
Qwen3-VL is a family of multimodal language models for interpreting images, video, and text. The repository provides inference examples, deployment guidance, and cookbooks for tasks such as OCR, spatial reasoning, and visual agents.

PYAS: Layered Endpoint Security for Windows
PYAS is a Windows endpoint security application that combines local machine-learning and YARA scanning with real-time monitoring, optional cloud analysis, and kernel-level controls. It suits security researchers and Windows users evaluating layered protection, with driver and remediation features best tested in an isolated environment.

feynman: Research Papers with an AI Agent
Feynman is a TypeScript CLI agent for investigating research topics, reviewing papers, and comparing sources. It suits researchers and developers who want structured, citation-aware literature work and can configure a model provider.

VoxCPM: Generate and Clone Multilingual Speech
VoxCPM is a tokenizer-free text-to-speech system for multilingual speech generation, voice design, and voice cloning. Its VoxCPM2 release targets teams and developers who need expressive speech synthesis and can support a 2B-parameter model.

MOSS-TTS: Generate Speech, Dialogue, and Sound with AI
MOSS-TTS is a family of speech and sound generation models for long-form narration, voice cloning, dialogue, voice design, sound effects, and streaming TTS. It offers multiple model architectures and deployment paths for research, production, and local inference.

autoresearch: Let AI Agents Run LLM Training Experiments
autoresearch gives an AI coding agent a small, single-GPU language-model training setup to modify and test autonomously. It suits researchers with supported NVIDIA hardware who want to explore training changes through repeated, time-bounded experiments.

ai-engineering-from-scratch: Learn AI by Building It
A free, MIT-licensed curriculum that teaches AI engineering through runnable lessons and hands-on artifacts. It covers foundations through LLMs, agents, protocols, and production, with paths for beginners and experienced developers.

FinceptTerminal: Research Markets with a Desktop Finance Terminal
FinceptTerminal is a native desktop application for financial research, combining market data, analytics, AI agents, and paper trading. It suits individual researchers and developers who can bring their own data and model credentials and accept the AGPL-3.0 license.

tensorrec: Build Custom Recommendation Systems with TensorFlow
TensorRec is a Python framework for building recommendation systems with TensorFlow. It combines user and item features with interaction data, while letting developers customize representation and loss functions. The project is not under active development.