Open Source Data Science Projects
Discover 42 open source Data Science repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Science projects here are most often combined with Python, Machine Learning and Library. Last updated October 3, 2026.
42 repositories · updated October 3, 2026

TabSTAR: Apply a Tabular Foundation Model to Data with Text Fields
TabSTAR is a Python model for classification and regression on tabular datasets that include text fields. Use its package to fit a pretrained model to your data, or its research tools to pretrain and evaluate on benchmarks.

cupy: Run NumPy and SciPy Workloads on GPUs
CuPy is a Python array library that brings NumPy- and SciPy-compatible computing to NVIDIA CUDA and AMD ROCm GPUs. It suits Python users who want GPU acceleration while reusing familiar APIs, or who need lower-level GPU controls.

graphic-walker: Embed Interactive Data Visualization in React
Graphic Walker is a TypeScript and React component for exploring datasets with drag-and-drop charts, filters, and data explanations. Embed it in an application when you need lightweight visual analytics rather than a full BI platform.

nolds: Nonlinear Measures for Dynamical Systems in Python
nolds is a Python library for calculating nonlinear measures in dynamical systems, specifically designed for one-dimensional time series. It provides implementations for various metrics such as sample entropy, correlation dimension, Lyapunov exponents, and Hurst exponent. This tool is valuable for analyzing the complexity, predictability, and memory of time series data, serving as both a practical utility and a learning resource.

scikit-learn: The Essential Python Library for Machine Learning
scikit-learn is a widely-used open-source Python library for machine learning, built upon SciPy. It provides a comprehensive suite of tools for data mining and data analysis, making it an indispensable resource for developers and data scientists. With its extensive algorithms and user-friendly interface, scikit-learn simplifies complex machine learning tasks.

gs-quant: Build Quantitative Finance Tools in Python
GS Quant is a Python toolkit for developing quantitative trading strategies, analyzing derivatives, and supporting risk management. It suits developers and quants with access to Goldman Sachs APIs who need finance-focused tools and analytics.

Box: Access Python Dictionaries with Dot Notation
Box is a Python library that adds attribute-style access to dictionaries while keeping key-based access available. It recursively converts nested dictionaries and lists, and includes helpers for converting data to common formats.

DataScienceInteractivePython: Learn Data Science Through Dashboards
Interactive Python dashboards let learners explore statistics, machine learning, and geostatistics by changing inputs and observing results. Designed for students and practitioners who benefit from hands-on experimentation.

multiresolution-time-series-transformer: Forecast Time Series at Multiple Scales
A PyTorch forecasting model that processes time series at several temporal resolutions, then fuses the representations to predict future values. It suits experimentation with multi-scale forecasting, with adaptations from the paper that users should account for.

lance: Store and Query Multimodal Lakehouse Data
Lance is a Rust-based lakehouse format and SDK for AI and machine-learning data. It combines columnar storage with random access, vector and full-text search, and dataset versioning for teams working with multimodal data.

superset: Explore Data and Build Business Intelligence Dashboards
Apache Superset is a web-based business intelligence platform for exploring SQL data, building visualizations, and creating dashboards. It suits analysts and organizations that want self-service analytics across supported data sources, with administrators managing deployment and access.

data-prep-kit: Prepare Data for LLM Applications
Data-Prep-Kit is a toolkit for cleaning, transforming, and enriching unstructured data used in LLM training and RAG pipelines. It offers reusable transforms that run with Python or Ray, from local experiments to larger-scale processing.