DataDreamer: Generate Synthetic Data and Train LLMs

Summary
DataDreamer is a Python library for building LLM workflows, generating synthetic datasets, and training or aligning models. It suits researchers and developers who want reproducible, resumable workflows across open-source and API-based models.
At a glance
- Language
- Python
- License
- MIT
- Stars
- 1.1k
- Forks
- 59
- Added to OSRepos
- July 3, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
DataDreamer brings LLM prompting, synthetic-data generation, and model training into a Python workflow. It addresses the work of connecting these stages and making experiments easier to resume, reproduce, and share.
It is aimed at researchers and developers who need to build multi-step LLM workflows or create training data, then fine-tune, instruction-tune, distill, or align models. It supports open-source and API-based LLMs, so it can fit both model experimentation and data-preparation pipelines.
Key Features
- Build multi-step prompting workflows with open-source or API-based LLMs.
- Generate synthetic datasets for new tasks or augment existing datasets.
- Fine-tune, instruction-tune, distill, and align models using existing or synthetic data.
- Cache work and resume workflows to support more efficient iteration.
- Use quantization and parameter-efficient training techniques such as LoRA.
- Share workflows and publish datasets and models with generated data cards, model cards, and citation information.
Use Cases
- Researchers can run reproducible experiments that connect LLM prompting, data generation, and model training.
- ML engineers can create synthetic examples to bootstrap a dataset or augment existing training data.
- Teams can build instruction-tuning or distillation pipelines when they have data and a suitable model to train.
- Developers can prototype workflows that combine API-based and open-source LLMs.
Project Facts
- Language: Python
- License: MIT
- Stars: 1.1k
- Forks: 59
- Topics: alignment, deep-learning, fine-tuning, gpt, instruction-tuning, llm, llmops, llms, machine-learning, natural-language-processing, nlp, nlp-library, openai, python, pytorch, synthetic-data, synthetic-dataset-generation, transformers
- Archived: no
Getting Started
Install the package:
pip3 install datadreamer.dev
See the README and documentation for setup details and examples.
Alternatives
- LLMBox: LLMBox centers on unified LLM training and evaluation, while DataDreamer focuses on reproducible workflows for data generation, training, and alignment.
- LlamaFactory: LlamaFactory focuses on fine-tuning through CLI and web interfaces, while DataDreamer also provides workflow tools for synthetic data generation and research.
- torchtune: torchtune provides editable PyTorch recipes for model post-training, while DataDreamer offers broader resumable workflows for data generation and model training.
- RL4LMs: RL4LMs specializes in reinforcement learning from custom rewards, while DataDreamer supports broader LLM workflows, including synthetic data and multiple training approaches.
Considerations
- Training and alignment workflows require appropriate model and compute resources; the supplied project information does not specify hardware requirements.
- The library supports several workflow stages, so users should consult the documentation for compatible models, integrations, and configuration details.
- The repository is not archived. Its listed latest push was on 2025-02-02, which is a snapshot and does not establish current maintenance activity.
Comparisons
Source repository
Open the original repository on GitHub.
26 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

agentevals: Evaluate AI Agents from OpenTelemetry Traces
October 4, 2026
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.