instructor: Extract Validated Structured Data from LLMs

Summary
Instructor is a Python library that turns LLM responses into validated, typed data using Pydantic models. It suits applications that need reliable extraction across supported model providers without writing custom parsing and validation flows.
At a glance
- Language
- Python
- License
- MIT
- Stars
- 14k
- Forks
- 1.3k
- Added to OSRepos
- October 12, 2025
- Last analyzed
- October 3, 2026
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Instructor is a Python library for requesting structured responses from language models and validating them against Pydantic models. It helps developers turn unstructured text into application-ready data while reducing the custom schema handling, parsing, and validation code needed for that workflow.
It is a good fit when structured extraction is the main requirement and a schema-first interface is preferable to building a full agent system. The project documents support for multiple model providers, with a common API for creating requests.
Key Features
- Define expected responses as Pydantic models for typed, validated results.
- Retry requests when generated data fails validation.
- Stream partial model instances as responses are generated.
- Represent nested data with nested models.
- Use a shared interface across providers listed in the README, including OpenAI, Anthropic, Google, and Ollama.
- Pass API keys directly or configure them through the environment.
Use Cases
- Extract names, dates, or other fields from customer messages for application workflows.
- Convert product descriptions or documents into typed records that downstream code can validate.
- Build intake forms or enrichment pipelines where invalid model output should be retried.
- Display progressively extracted fields in an interface using streamed partial results.
Project Facts
- Language: Python
- License: MIT
- Stars: 14k
- Forks: 1.3k
- Topics: openai, openai-function-calli, openai-functions, pydantic-v2, python, validation
- Archived: no
Getting Started
Install with:
pip install instructor
See the README and Python documentation for provider setup and usage details.
Considerations
- This library relies on language-model providers, so using a hosted provider may require an API key and incur provider costs. Ollama is listed as a local option.
- Validation and retries can improve the shape of returned data, but do not guarantee that extracted values are factually correct.
- The README positions Instructor around schema-first extraction, not richer agent workflows. Choose a broader agent runtime if that is the central need.
Source repository
Open the original repository on GitHub.
23 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

agentevals: Evaluate AI Agents from OpenTelemetry Traces
October 4, 2026
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.