Open Source Testing Projects
Discover 37 open source Testing repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Testing projects here are most often combined with Python, LLM and Developer Tools. Last updated October 3, 2026.
37 repositories · updated October 3, 2026

model_mommy: Create Test Fixtures for Python Projects
Model Mommy is a Python library for creating fixtures for tests. It is no longer maintained; Model Bakery is its replacement, and existing users are encouraged to migrate.

fake2db: Generate Custom Test Databases with Fake Data
fake2db is a powerful Python utility designed to create custom test databases populated with fake, yet valid, data. It supports a wide array of popular database systems, including SQLite, MySQL, PostgreSQL, MongoDB, Redis, and CouchDB. This tool is ideal for developers and testers needing quick, realistic data for testing and development environments.

faker: Generate Realistic Fake Data in Python
Faker is a Python library for generating synthetic names, addresses, text, and other data. Use it to populate test databases, create sample documents, or anonymize production data, with support for localized output and custom providers.

Mimesis: A Powerful Python Library for Realistic Fake Data Generation
Mimesis is a robust Python library designed for generating fake yet realistic data across various languages and locales. It simplifies the creation of diverse data types, from personal information to financial details. This makes it an invaluable tool for development, testing, and anonymization tasks.

werkzeug: Build WSGI Web Applications and Utilities
Werkzeug is a Python library for building WSGI applications, with HTTP request and response tools, routing, and testing utilities. It suits developers who want lower-level control over web application behavior without a framework imposing extra structure.

pgrust: Run a Rust Rewrite of PostgreSQL
pgrust is a Rust reimplementation of PostgreSQL that aims to preserve its wire protocol and SQL behavior while reworking execution, concurrency, and storage. It is for evaluation and experimentation today, not production data you cannot afford to lose.

evalplus: Rigorously Evaluate LLM-Generated Code
EvalPlus evaluates code generated by language models with expanded correctness tests for HumanEval and MBPP, plus efficiency checks through EvalPerf. It is for researchers and developers comparing models or validating generated code more rigorously.

agentevals: Evaluate AI Agent Execution Trajectories
AgentEvals provides Python and TypeScript evaluators for checking the steps AI agents take, including tool calls and graph paths. Use it to compare runs with references or have an LLM judge trajectory quality.

JailbreakEval: Compare LLM Jailbreak Evaluators
JailbreakEval brings together automated methods for assessing whether language-model responses comply with jailbreak attempts. Researchers can compare evaluators across datasets, while developers can build and benchmark new evaluation methods.

bruno: Explore and Test APIs with a Local-First Client
Bruno is an API client for exploring requests, testing APIs, and running collections from a desktop app or CLI. It stores collections as plain-text files on your device, making them suitable for Git-based collaboration without cloud sync.

promptfoo: Evaluate and Red-Team LLM Applications
Promptfoo is a CLI and library for evaluating prompts, comparing models, and testing LLM applications for security risks. It suits developers who want repeatable quality and vulnerability checks locally or in CI/CD.

GeoPort: Simulate an iPhone's Location from Desktop
GeoPort is a desktop app for simulating an iOS device's location on iOS 17 and 18. It is aimed at developers testing location-based apps and users who need to change a device's simulated location, with Windows and macOS packages available from releases.