Open Source Data Science Projects
Discover 46 open source Data Science repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Science projects here are most often combined with Python, Machine Learning and Library. Last updated October 3, 2026.
46 repositories · updated October 3, 2026

multiresolution-time-series-transformer: Forecast Time Series at Multiple Scales
A PyTorch forecasting model that processes time series at several temporal resolutions, then fuses the representations to predict future values. It suits experimentation with multi-scale forecasting, with adaptations from the paper that users should account for.

lance: Store and Query Multimodal Lakehouse Data
Lance is a Rust-based lakehouse format and SDK for AI and machine-learning data. It combines columnar storage with random access, vector and full-text search, and dataset versioning for teams working with multimodal data.

superset: Explore Data and Build Business Intelligence Dashboards
Apache Superset is a web-based business intelligence platform for exploring SQL data, building visualizations, and creating dashboards. It suits analysts and organizations that want self-service analytics across supported data sources, with administrators managing deployment and access.

data-prep-kit: Prepare Data for LLM Applications
Data-Prep-Kit is a toolkit for cleaning, transforming, and enriching unstructured data used in LLM training and RAG pipelines. It offers reusable transforms that run with Python or Ray, from local experiments to larger-scale processing.

gradio: Build and Share Python Web Apps for Machine Learning
Gradio turns Python functions and machine-learning models into interactive web apps without requiring frontend development. Use it to prototype interfaces, share demos, or build more customized apps with components and event-driven layouts.

QGIS: Create, Analyze, and Share Geographic Maps
QGIS is a cross-platform GIS for working with spatial data, creating maps, and running geospatial analysis. It suits GIS professionals, researchers, public agencies, and other users who need a flexible desktop mapping and data workflow.

numba: Compile Numerical Python to Machine Code
Numba compiles numerically focused Python functions to machine code using LLVM, with support for many NumPy operations. It suits Python developers who need faster numerical workloads, loop parallelization, or GPU code generation without rewriting everything in a lower-level language.

Flyte: Scalable Workflow Orchestration for Data and ML
Flyte is an open-source, scalable, and flexible workflow orchestration platform that seamlessly unifies data, machine learning, and analytics stacks. It leverages Kubernetes as its underlying platform, enabling the construction of robust and reproducible production-grade pipelines.

plexe: Build Machine Learning Models from Prompts
Plexe turns a natural-language description and tabular dataset into a trained, packaged machine learning model through a multi-agent workflow. It is aimed at developers and data teams who want to automate model selection and iteration while retaining control over configuration and deployment.

FinRL-Trading: Build and Deploy Quantitative Trading Strategies
FinRL-Trading is a modular Python system for developing, backtesting, and executing quantitative trading strategies. Its shared portfolio-weight interface connects strategy components to backtests and Alpaca trading, making it useful for researchers and practitioners evaluating end-to-end workflows.