Open Source Machine Learning Projects
Discover 191 open source Machine Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Machine Learning projects here are most often combined with Python, Deep Learning and LLM. Last updated October 4, 2026.
191 repositories · updated October 4, 2026

xTuring: Fine-Tune and Run Personalized LLMs
xTuring is a Python library for preparing data, fine-tuning, evaluating, and running open-source language models locally or in a private cloud. It is aimed at developers who want model customization through a high-level API and methods such as LoRA and quantization.

RL4LMs: Fine-Tune Language Models with Reinforcement Learning
RL4LMs is a Python library for training language models against custom reward functions using on-policy reinforcement learning. It suits NLP researchers and developers who need configurable training components for text-generation tasks.

torchtune: Fine-Tune and Post-Train Large Language Models
torchtune is a PyTorch library for configuring and running LLM post-training workflows, from supervised fine-tuning to preference optimization. It is aimed at developers who want editable recipes and model-specific configs, but is no longer actively maintained.

RouteLLM: Route LLM Requests to Balance Cost and Quality
RouteLLM is a Python framework for routing prompts between stronger, costlier models and cheaper alternatives. It helps teams manage the cost-quality tradeoff, serve requests through an OpenAI-compatible interface, and evaluate routing strategies.

RAGChecker: Diagnose Retrieval-Augmented Generation Systems
RAGChecker evaluates RAG pipelines with overall, retriever, and generator metrics. It is for developers and researchers who need to identify whether retrieval or generation is driving quality problems and guide targeted improvements.

rerankers: Use Diverse Reranking Models Through One Python API
rerankers provides a shared Python interface for reranking documents with cross-encoders, LLM-based methods, and hosted APIs. It suits developers building retrieval systems who want to compare or switch rerankers without adapting their application to each model's interface.

llm-compressor: Compress Models for vLLM Inference
LLM Compressor applies post-training quantization and related model transformations to prepare Hugging Face models for vLLM deployment. It suits teams seeking smaller or more inference-efficient checkpoints, with support for multiple formats and large-model workflows.

LightLLM: Serve Large Language Models with a Python Framework
LightLLM is a Python framework for LLM inference and serving, designed for scalable deployment and fast generation. It suits teams operating model-serving systems and researchers building on inference components.

torchchat: Run PyTorch LLMs on Desktop, Server, and Mobile
torchchat is a PyTorch codebase for running and interacting with language models locally through Python, native C++ runners, and mobile apps. It supports several execution and export paths, but is no longer under active development.

DataDreamer: Generate Synthetic Data and Train LLMs
DataDreamer is a Python library for building LLM workflows, generating synthetic datasets, and training or aligning models. It suits researchers and developers who want reproducible, resumable workflows across open-source and API-based models.

deepfabric: Generate and Evaluate Synthetic Training Data
DeepFabric generates domain-specific synthetic datasets for language-model training and agent evaluation. It combines topic planning, tool-use traces, validation, and evaluation in a Python library and CLI.

EasyInstruct: Generate, Select, and Prompt LLM Instructions
EasyInstruct is a Python framework for preparing instruction data and prompts for large language model research. It combines instruction generation and dataset selection tools with prompt and local-model execution modules.