DataDreamer: Generate Synthetic Data and Train LLMs

DataDreamer: Generate Synthetic Data and Train LLMs

Summary

DataDreamer is a Python library for building LLM workflows, generating synthetic datasets, and training or aligning models. It suits researchers and developers who want reproducible, resumable workflows across open-source and API-based models.

At a glance

Language
Python
License
MIT
Stars
1.1k
Forks
59
Added to OSRepos
July 3, 2026
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

DataDreamer brings LLM prompting, synthetic-data generation, and model training into a Python workflow. It addresses the work of connecting these stages and making experiments easier to resume, reproduce, and share.

It is aimed at researchers and developers who need to build multi-step LLM workflows or create training data, then fine-tune, instruction-tune, distill, or align models. It supports open-source and API-based LLMs, so it can fit both model experimentation and data-preparation pipelines.

Key Features

  • Build multi-step prompting workflows with open-source or API-based LLMs.
  • Generate synthetic datasets for new tasks or augment existing datasets.
  • Fine-tune, instruction-tune, distill, and align models using existing or synthetic data.
  • Cache work and resume workflows to support more efficient iteration.
  • Use quantization and parameter-efficient training techniques such as LoRA.
  • Share workflows and publish datasets and models with generated data cards, model cards, and citation information.

Use Cases

  • Researchers can run reproducible experiments that connect LLM prompting, data generation, and model training.
  • ML engineers can create synthetic examples to bootstrap a dataset or augment existing training data.
  • Teams can build instruction-tuning or distillation pipelines when they have data and a suitable model to train.
  • Developers can prototype workflows that combine API-based and open-source LLMs.

Project Facts

  • Language: Python
  • License: MIT
  • Stars: 1.1k
  • Forks: 59
  • Topics: alignment, deep-learning, fine-tuning, gpt, instruction-tuning, llm, llmops, llms, machine-learning, natural-language-processing, nlp, nlp-library, openai, python, pytorch, synthetic-data, synthetic-dataset-generation, transformers
  • Archived: no

Getting Started

Install the package:

pip3 install datadreamer.dev

See the README and documentation for setup details and examples.

Alternatives

  • LLMBox: LLMBox centers on unified LLM training and evaluation, while DataDreamer focuses on reproducible workflows for data generation, training, and alignment.
  • LlamaFactory: LlamaFactory focuses on fine-tuning through CLI and web interfaces, while DataDreamer also provides workflow tools for synthetic data generation and research.
  • torchtune: torchtune provides editable PyTorch recipes for model post-training, while DataDreamer offers broader resumable workflows for data generation and model training.
  • RL4LMs: RL4LMs specializes in reinforcement learning from custom rewards, while DataDreamer supports broader LLM workflows, including synthetic data and multiple training approaches.

Considerations

  • Training and alignment workflows require appropriate model and compute resources; the supplied project information does not specify hardware requirements.
  • The library supports several workflow stages, so users should consult the documentation for compatible models, integrations, and configuration details.
  • The repository is not archived. Its listed latest push was on 2025-02-02, which is a snapshot and does not establish current maintenance activity.

Source repository

Open the original repository on GitHub.

25 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design

web-design: A Claude Code SKILL for Spec-First Web Page Design

October 3, 2026

The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

Claude CodeClaude SkillDesign System
OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner

OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner

October 2, 2026

OOMWOO is an ambitious open-source project enabling users to build their own robot vacuum cleaner using Raspberry Pi, 3D printing, and ROS2. It emphasizes local operation, hackability, and integration with Home Assistant, providing a high-quality, customizable home appliance. This project aims to deliver a fully open hardware, software, and firmware solution for autonomous home cleaning.

RoboticsOpen SourceDiy
Shepherd: Reversible Execution Traces for Programmable Meta-Agents

Shepherd: Reversible Execution Traces for Programmable Meta-Agents

October 2, 2026

Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

PythonAIAgent Framework
Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

October 1, 2026

Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

PythonAIAgent
OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️