Open Source Machine Learning Projects
Discover 191 open source Machine Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Machine Learning projects here are most often combined with Python, Deep Learning and LLM. Last updated October 4, 2026.
191 repositories · updated October 4, 2026

lighteval: Evaluate Language Models Across Backends
Lighteval is a Python toolkit for running LLM evaluations across local models and remote inference backends. It combines a broad task catalog with custom metrics and detailed sample-level results for teams comparing or debugging model performance.

promptbench: Evaluate LLMs and Test Prompt Robustness
PromptBench is a Python library for evaluating language and multimodal models across datasets, prompting methods, and adversarial attacks. It suits researchers and developers comparing model behavior or studying robustness and dynamic evaluation.

evidently: Evaluate and Monitor ML and LLM Systems
Evidently is a Python framework for evaluating, testing, and monitoring machine-learning and LLM systems, from data quality to generated text. Use it to build offline reports and regression checks or track metrics over time in a monitoring dashboard.

xgrammar: Constrain Language Model Output to Structured Formats
XGrammar is a library for constrained decoding that helps language models produce outputs matching JSON, regular expressions, or context-free grammars. It suits teams integrating structured output into LLM inference and applications that need reliable machine-readable responses.

jsonformer: Generate Schema-Conforming JSON with Language Models
Jsonformer guides Hugging Face language models to produce JSON that matches a supplied schema by generating variable content while inserting predictable structure itself. It suits developers who need structured model output and can work within its supported JSON Schema subset.

JailbreakEval: Compare LLM Jailbreak Evaluators
JailbreakEval brings together automated methods for assessing whether language-model responses comply with jailbreak attempts. Researchers can compare evaluators across datasets, while developers can build and benchmark new evaluation methods.

EasyJailbreak: Build and Evaluate LLM Jailbreak Attacks
EasyJailbreak is a Python framework for assembling and testing jailbreak methods against language models. It suits researchers and developers who need reusable attack components and a structured way to evaluate model responses.

hiring-agent: Score Resumes with AI and GitHub Signals
Hiring Agent turns PDF resumes into structured profiles, enriches them with GitHub data, and scores them against configurable role rubrics. It is for teams exploring explainable resume prioritization, with human review remaining essential.

spacy-llm: Add LLM Tasks to spaCy NLP Pipelines
spacy-llm connects large language models to spaCy pipelines, turning model responses into structured NLP outputs without training data. It suits teams prototyping NLP tasks or combining LLM components with conventional spaCy processing.

MarkLLM: Implement and Evaluate LLM Watermarking
MarkLLM is a Python toolkit for implementing, visualizing, and evaluating text watermarking methods for large language models. It helps researchers compare watermark detection, robustness, and text-quality effects through shared APIs and evaluation pipelines.

easy-whisper-ui: Transcribe Audio and Video Locally
EasyWhisperUI is a desktop app for transcribing audio and video locally with whisper.cpp. It offers batch and live transcription, subtitle and text exports, and GPU acceleration where supported, without requiring command-line use.

REAL-Video-Enhancer: Interpolate and Upscale Videos
A desktop video-processing app for interpolating frames, upscaling, decompressing, and denoising video on Windows, Linux, and macOS. It offers several inference backends for different GPU hardware.