Open Source Machine Learning Projects
Discover 191 open source Machine Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Machine Learning projects here are most often combined with Python, Deep Learning and LLM. Last updated October 4, 2026.
191 repositories · updated October 4, 2026

JARVIS: Connect Language Models to AI Tools
JARVIS, also known as HuggingGPT, uses an LLM to plan requests, choose expert models from Hugging Face, run tasks, and combine their results. It suits research and prototyping that explores natural-language orchestration of AI models.

peft: Fine-Tune Models with Fewer Trainable Parameters
Hugging Face PEFT adapts pretrained models by training a small set of additional parameters instead of updating the full model. It is for practitioners who want to reduce fine-tuning compute and storage costs across supported model workflows.

fastFM: Build Factorization Machine Models in Python
fastFM is a Python library for factorization machines, with scikit-learn-style estimators for regression, classification, and ranking. It suits machine-learning practitioners working with sparse feature data and recommendation problems.

llmgym: Build and Benchmark LLM Applications That Learn
LLMGym provides a shared environment interface for developing and evaluating LLM applications that learn from feedback. It is aimed at researchers and developers who want to test agents across varied interactive tasks.

tortoise.cpp: Generate Speech with Tortoise TTS in C++
tortoise.cpp is a C++ implementation of Tortoise TTS built on ggml for local text-to-speech generation. It supports CPU and CUDA builds, with Metal support described as in progress.

Qwen3: Build and Run Open-Weight Language Models
Qwen3 is a family of open-weight language models for reasoning, chat, coding, and tool use. This repository guides developers through model selection, local inference, deployment, and fine-tuning with supported frameworks.

AI-Scientist-v2: Automate Scientific Experiments and Paper Drafts
AI-Scientist-v2 uses LLM agents and tree search to propose research ideas, run experiments, analyze results, and draft papers. It is aimed at researchers exploring open-ended machine learning questions who can provide model APIs and a controlled compute environment.

pytorch-image-models: Use PyTorch Image Models and Training Tools
A Python library of PyTorch image models, pretrained weights, training utilities, and reference scripts. Use it to build image-classification systems, extract backbone features, or train and evaluate models across a broad range of architectures.

kapre: Add Audio Preprocessing Layers to Keras Models
Kapre provides TensorFlow-backed Keras layers for transforming audio waveforms into STFT, mel-spectrogram, and related features inside a model. It suits audio ML developers who want preprocessing to travel with training and deployment.

WeClone: Create an AI Avatar from Chat History
WeClone turns exported chat histories into fine-tuned language models that can respond in a person's style. It combines data preparation, privacy filtering, model training, and chatbot deployment for people comfortable running local AI workflows.

textgen: Run Local Language Models in a Desktop App
TextGen is a private desktop and web interface for running local language models, with support for chat, vision, tools, and compatible APIs. It suits people who want to use and manage models on their own hardware.

Qwen: Run and Fine-Tune Pretrained Language Models
Qwen is Alibaba Cloud’s Python repository for running and adapting pretrained and chat language models, with an emphasis on Chinese and English. It includes inference, quantization, fine-tuning, and deployment paths, but is no longer actively maintained; the project points users to Qwen2.