Open Source Python Projects
Discover 606 open source Python repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Python projects here are most often combined with Library, Developer Tools and LLM. Last updated October 4, 2026.
606 repositories · updated October 4, 2026

Agentarium: Build and Orchestrate AI Agent Simulations
Agentarium is a Python framework for creating AI agents that interact, take context-based actions, and retain memories. Use it to prototype multi-agent scenarios, add custom actions, and save agent states for repeatable experiments.

lighteval: Evaluate Language Models Across Backends
Lighteval is a Python toolkit for running LLM evaluations across local models and remote inference backends. It combines a broad task catalog with custom metrics and detailed sample-level results for teams comparing or debugging model performance.

promptbench: Evaluate LLMs and Test Prompt Robustness
PromptBench is a Python library for evaluating language and multimodal models across datasets, prompting methods, and adversarial attacks. It suits researchers and developers comparing model behavior or studying robustness and dynamic evaluation.

langtest: Test Language Models for Safety and Quality
LangTest is a Python library for generating and running tests that assess language-model quality, including robustness, bias, fairness, and accuracy. It helps NLP and AI teams identify issues and, for select models, augment training data based on results.

evalplus: Rigorously Evaluate LLM-Generated Code
EvalPlus evaluates code generated by language models with expanded correctness tests for HumanEval and MBPP, plus efficiency checks through EvalPerf. It is for researchers and developers comparing models or validating generated code more rigorously.

agentevals: Evaluate AI Agent Execution Trajectories
AgentEvals provides Python and TypeScript evaluators for checking the steps AI agents take, including tool calls and graph paths. Use it to compare runs with references or have an LLM judge trajectory quality.

evidently: Evaluate and Monitor ML and LLM Systems
Evidently is a Python framework for evaluating, testing, and monitoring machine-learning and LLM systems, from data quality to generated text. Use it to build offline reports and regression checks or track metrics over time in a monitoring dashboard.

OpenMontage: Produce Videos with AI Coding Assistants
OpenMontage gives AI coding assistants structured pipelines for researching, scripting, generating assets, editing, and rendering videos. It suits creators and developers who want to automate production while retaining control over choices and approvals.

phoenix: Observe and Evaluate AI Applications
Arize Phoenix is a self-hosted platform for tracing, evaluating, and troubleshooting LLM applications. It helps AI engineers inspect runtime behavior and test changes to prompts, models, and retrieval.

observers: Track and Store AI API Interactions
Observers wraps generative AI clients to capture interactions and sync them to storage backends. It suits Python teams that need lightweight observability across supported LLM providers, with storage options ranging from DuckDB to OpenTelemetry-compatible services.

jsonformer: Generate Schema-Conforming JSON with Language Models
Jsonformer guides Hugging Face language models to produce JSON that matches a supplied schema by generating variable content while inserting predictable structure itself. It suits developers who need structured model output and can work within its supported JSON Schema subset.

JailbreakEval: Compare LLM Jailbreak Evaluators
JailbreakEval brings together automated methods for assessing whether language-model responses comply with jailbreak attempts. Researchers can compare evaluators across datasets, while developers can build and benchmark new evaluation methods.