Open Source Data Pipelines

Data pipelines move information through a sequence of steps, such as collecting, validating, transforming, and delivering data. They help teams automate repeatable workflows, connect systems with different formats, and keep data timely and consistent for analytics, applications, and machine learning. Pipelines can process records in batches or continuously as new data arrives, and may include tasks such as document extraction, filtering, and quality checks.

Open source tools in this area include workflow orchestrators, data processing libraries, connectors, and systems for monitoring runs and managing dependencies. When choosing one, consider its maturity, license, maintenance activity, deployment requirements, scaling model, and integration with your existing storage and compute platforms. These tools are useful for data engineers, researchers, and application teams building reliable workflows, from small automated tasks to large-scale processing systems.

2 repositories · updated December 15, 2025

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️