Open Source Data Pipelines
Data pipelines move information through a sequence of steps, such as collecting, validating, transforming, and delivering data. They help teams automate repeatable workflows, connect systems with different formats, and keep data timely and consistent for analytics, applications, and machine learning. Pipelines can process records in batches or continuously as new data arrives, and may include tasks such as document extraction, filtering, and quality checks.
Open source tools in this area include workflow orchestrators, data processing libraries, connectors, and systems for monitoring runs and managing dependencies. When choosing one, consider its maturity, license, maintenance activity, deployment requirements, scaling model, and integration with your existing storage and compute platforms. These tools are useful for data engineers, researchers, and application teams building reliable workflows, from small automated tasks to large-scale processing systems.
2 repositories · updated December 15, 2025

dlt: The Open-Source Python Library for Easy Data Loading
dlt, the data load tool, is an open-source Python library designed to simplify and automate data loading tasks. It efficiently extracts, normalizes, and loads data from various sources into well-structured datasets. Highly versatile, dlt supports diverse data sources and destinations, making it suitable for deployment in a wide range of environments.

Flyte: Scalable Workflow Orchestration for Data and ML
Flyte is an open-source, scalable, and flexible workflow orchestration platform that seamlessly unifies data, machine learning, and analytics stacks. It leverages Kubernetes as its underlying platform, enabling the construction of robust and reproducible production-grade pipelines.