instructor: Extract Validated Structured Data from LLMs

instructor: Extract Validated Structured Data from LLMs

Summary

Instructor is a Python library that turns LLM responses into validated, typed data using Pydantic models. It suits applications that need reliable extraction across supported model providers without writing custom parsing and validation flows.

At a glance

Language
Python
License
MIT
Stars
14k
Forks
1.3k
Added to OSRepos
October 12, 2025
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

Instructor is a Python library for requesting structured responses from language models and validating them against Pydantic models. It helps developers turn unstructured text into application-ready data while reducing the custom schema handling, parsing, and validation code needed for that workflow.

It is a good fit when structured extraction is the main requirement and a schema-first interface is preferable to building a full agent system. The project documents support for multiple model providers, with a common API for creating requests.

Key Features

  • Define expected responses as Pydantic models for typed, validated results.
  • Retry requests when generated data fails validation.
  • Stream partial model instances as responses are generated.
  • Represent nested data with nested models.
  • Use a shared interface across providers listed in the README, including OpenAI, Anthropic, Google, and Ollama.
  • Pass API keys directly or configure them through the environment.

Use Cases

  • Extract names, dates, or other fields from customer messages for application workflows.
  • Convert product descriptions or documents into typed records that downstream code can validate.
  • Build intake forms or enrichment pipelines where invalid model output should be retried.
  • Display progressively extracted fields in an interface using streamed partial results.

Project Facts

  • Language: Python
  • License: MIT
  • Stars: 14k
  • Forks: 1.3k
  • Topics: openai, openai-function-calli, openai-functions, pydantic-v2, python, validation
  • Archived: no

Getting Started

Install with:

pip install instructor

See the README and Python documentation for provider setup and usage details.

Considerations

  • This library relies on language-model providers, so using a hosted provider may require an API key and incur provider costs. Ollama is listed as a local option.
  • Validation and retries can improve the shape of returned data, but do not guarantee that extracted values are factually correct.
  • The README positions Instructor around schema-first extraction, not richer agent workflows. Choose a broader agent runtime if that is the central need.

Source repository

Open the original repository on GitHub.

23 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️