Open Source Machine Learning Projects
Discover 191 open source Machine Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Machine Learning projects here are most often combined with Python, Deep Learning and LLM. Last updated October 4, 2026.
191 repositories · updated October 4, 2026

aimet: Quantize and Compress Trained Neural Networks
AIMET helps ML engineers reduce the compute and memory needs of trained neural networks through quantization and compression. It supports PyTorch and ONNX workflows for preparing models for efficient deployment, including on edge devices.

llm-reasoners: Build and Inspect LLM Reasoning Algorithms
LLM Reasoners is a Python library for building and evaluating multi-step reasoning methods with large language models. It suits researchers and developers who need reusable search algorithms, model backends, and tools to inspect reasoning traces.

awesome-AI-books: Find AI Books, Papers, and Learning Resources
A curated directory of books, papers, online references, and practice environments across AI and related fields. It helps learners find starting points for topics from mathematics and machine learning to reinforcement learning and quantum computing.

parakeet-mlx: Transcribe Audio on Apple Silicon
parakeet-mlx runs Nvidia Parakeet speech recognition models on Apple Silicon using MLX. It provides command-line and Python interfaces for transcription, timestamps, chunking, and streaming audio.

faiss: Search and Cluster Dense Vectors Efficiently
Faiss provides indexing methods for similarity search and clustering over dense vectors, with trade-offs between search speed, accuracy, and memory use. It suits teams building vector retrieval or large-scale nearest-neighbor systems in C++ or Python.

datatrove: Build Large-Scale Text Data Pipelines
DataTrove is a Python library for processing, filtering, and deduplicating text datasets at scale. It combines reusable pipeline blocks with local and cluster executors, making it useful for data preparation workflows such as building LLM training corpora.

EasyEdit: Edit Knowledge and Steer Large Language Models
EasyEdit is a framework for changing specific knowledge or behavior in large language models and evaluating the effects. It brings together multiple editing and inference-time steering methods for researchers and developers comparing approaches or testing targeted edits.

cactus: Run AI Inference on Phones and Wearables
Cactus is a C++ inference engine for running language, vision, and speech models on mobile and edge devices. It combines quantization, device-focused kernels, and optional cloud handoff for applications that need local inference with a fallback for harder queries.

RecDebiasing: Find Research on Recommendation Bias
RecDebiasing is a curated research index for recommendation-system debiasing, with papers, datasets, and links to available code. It is useful to researchers and practitioners surveying methods for particular sources of bias.

vllm-cli: Manage and Serve LLMs with vLLM
vllm-cli provides interactive and command-line workflows for configuring and running vLLM model servers. It is aimed at developers and self-hosters who want model discovery, reusable profiles, and server monitoring in one terminal tool.

lmql: Program LLMs with Python-Like Constraints
LMQL combines Python-style control flow with language-model queries and constraints on generated text. It suits developers building structured, model-driven workflows who need more control than prompt templates provide.

Biomni: Run Biomedical Research Tasks with an AI Agent
Biomni is a Python agent for carrying out biomedical research tasks using language-model reasoning, retrieval, and code execution. It is aimed at researchers who want to connect natural-language questions with biomedical tools and data.