Open Source LLM Projects

Discover 273 open source LLM repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. LLM projects here are most often combined with Python, AI Agents and AI. Last updated October 4, 2026.

273 repositories · updated October 4, 2026

promptbench: Evaluate LLMs and Test Prompt Robustness

promptbench: Evaluate LLMs and Test Prompt Robustness

PromptBench is a Python library for evaluating language and multimodal models across datasets, prompting methods, and adversarial attacks. It suits researchers and developers comparing model behavior or studying robustness and dynamic evaluation.

PythonLLMMachine Learning
Added Jul 1, 2026 View details
langtest: Test Language Models for Safety and Quality

langtest: Test Language Models for Safety and Quality

LangTest is a Python library for generating and running tests that assess language-model quality, including robustness, bias, fairness, and accuracy. It helps NLP and AI teams identify issues and, for select models, augment training data based on results.

PythonLLMTesting
Added Jun 30, 2026 View details
evalplus: Rigorously Evaluate LLM-Generated Code

evalplus: Rigorously Evaluate LLM-Generated Code

EvalPlus evaluates code generated by language models with expanded correctness tests for HumanEval and MBPP, plus efficiency checks through EvalPerf. It is for researchers and developers comparing models or validating generated code more rigorously.

PythonLLMEvaluation
Added Jun 30, 2026 View details
agentevals: Evaluate AI Agent Execution Trajectories

agentevals: Evaluate AI Agent Execution Trajectories

AgentEvals provides Python and TypeScript evaluators for checking the steps AI agents take, including tool calls and graph paths. Use it to compare runs with references or have an LLM judge trajectory quality.

PythonTypeScriptAI Agents
Added Jun 30, 2026 View details
evidently: Evaluate and Monitor ML and LLM Systems

evidently: Evaluate and Monitor ML and LLM Systems

Evidently is a Python framework for evaluating, testing, and monitoring machine-learning and LLM systems, from data quality to generated text. Use it to build offline reports and regression checks or track metrics over time in a monitoring dashboard.

PythonMachine LearningLLM
Added Jun 30, 2026 View details
phoenix: Observe and Evaluate AI Applications

phoenix: Observe and Evaluate AI Applications

Arize Phoenix is a self-hosted platform for tracing, evaluating, and troubleshooting LLM applications. It helps AI engineers inspect runtime behavior and test changes to prompts, models, and retrieval.

PythonAILLM
Added Jun 28, 2026 View details
observers: Track and Store AI API Interactions

observers: Track and Store AI API Interactions

Observers wraps generative AI clients to capture interactions and sync them to storage backends. It suits Python teams that need lightweight observability across supported LLM providers, with storage options ranging from DuckDB to OpenTelemetry-compatible services.

PythonAILLM
Added Jun 28, 2026 View details
xgrammar: Constrain Language Model Output to Structured Formats

xgrammar: Constrain Language Model Output to Structured Formats

XGrammar is a library for constrained decoding that helps language models produce outputs matching JSON, regular expressions, or context-free grammars. It suits teams integrating structured output into LLM inference and applications that need reliable machine-readable responses.

C++LLMMachine Learning
Added Jun 27, 2026 View details
freellmapi: Route LLM Requests Through One API

freellmapi: Route LLM Requests Through One API

FreeLLMAPI is a self-hosted TypeScript router that combines free LLM provider accounts and custom OpenAI-compatible endpoints behind one API. It selects models, tracks quotas, and retries across providers, mainly for personal experimentation and development.

TypeScriptLLMAI
Added Jun 27, 2026 View details
jsonformer: Generate Schema-Conforming JSON with Language Models

jsonformer: Generate Schema-Conforming JSON with Language Models

Jsonformer guides Hugging Face language models to produce JSON that matches a supplied schema by generating variable content while inserting predictable structure itself. It suits developers who need structured model output and can work within its supported JSON Schema subset.

PythonLLMMachine Learning
Added Jun 27, 2026 View details
JailbreakEval: Compare LLM Jailbreak Evaluators

JailbreakEval: Compare LLM Jailbreak Evaluators

JailbreakEval brings together automated methods for assessing whether language-model responses comply with jailbreak attempts. Researchers can compare evaluators across datasets, while developers can build and benchmark new evaluation methods.

PythonLLMMachine Learning
Added Jun 26, 2026 View details
EasyJailbreak: Build and Evaluate LLM Jailbreak Attacks

EasyJailbreak: Build and Evaluate LLM Jailbreak Attacks

EasyJailbreak is a Python framework for assembling and testing jailbreak methods against language models. It suits researchers and developers who need reusable attack components and a structured way to evaluate model responses.

PythonLLMSecurity
Added Jun 26, 2026 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️