Open Source Data Science Projects
Discover 36 open source Data Science repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Science projects here are most often combined with Python, Machine Learning and Library. Last updated October 3, 2026.
36 repositories · updated October 3, 2026

pyAudioAnalysis: Extract Features and Analyze Audio in Python
pyAudioAnalysis is a Python library for extracting audio features and building classification, detection, and segmentation workflows. It suits researchers and developers who want an established toolkit for analyzing audio files with machine-learning methods.

jax: Transform and Accelerate Python Numerical Programs
JAX is a Python library for transforming numerical programs with automatic differentiation, compilation, and vectorization. Use it for high-performance scientific computing and machine learning, especially when workloads need to scale across accelerators.

ML-From-Scratch: Learn Machine Learning Through NumPy Implementations
ML-From-Scratch provides transparent Python and NumPy implementations of machine learning algorithms, from regression and clustering to neural networks and reinforcement learning. It is suited to learners who want to inspect how models work rather than use an optimized production framework.

rio: Build Web Apps in Python
Rio is a Python framework for building interactive web apps and local GUI applications with component-based interfaces. It suits Python developers who want to create polished interfaces without writing HTML, CSS, or JavaScript.

Spotlight: Deep Recommender Models with PyTorch
Spotlight is a Python library built on PyTorch for developing deep and shallow recommender models. It offers a comprehensive set of building blocks for various loss functions, representations, and utilities for handling recommendation datasets. This tool is designed for rapid exploration and prototyping of new recommender systems.

awesome-quant: Find Quantitative Finance Tools and Resources
awesome-quant is a categorized directory of quantitative finance libraries, data sources, trading systems, and learning materials. Use it to explore options across languages and research workflows, then evaluate each project for your needs.

datatrove: Build Large-Scale Text Data Pipelines
DataTrove is a Python library for processing, filtering, and deduplicating text datasets at scale. It combines reusable pipeline blocks with local and cluster executors, making it useful for data preparation workflows such as building LLM training corpora.

RecDebiasing: A Comprehensive Collection of Recommendation Debiasing Methods
RecDebiasing is a valuable GitHub repository that curates a wide array of debiasing methods for recommendation systems. It compiles recent research papers, relevant datasets, and associated codebases, offering a centralized resource for understanding and addressing various biases. This collection is essential for researchers and practitioners focused on building more fair and accurate recommender systems.

pyqtgraph: Build Fast Scientific Data Visualizations
PyQtGraph is a Python graphics library for interactive visualization in scientific and engineering applications. It combines NumPy, Qt and optional OpenGL support for plotting and displaying data in desktop applications.

Biomni: Run Biomedical Research Tasks with an AI Agent
Biomni is a Python agent for carrying out biomedical research tasks using language-model reasoning, retrieval, and code execution. It is aimed at researchers who want to connect natural-language questions with biomedical tools and data.

TabSTAR: Apply a Tabular Foundation Model to Data with Text Fields
TabSTAR is a Python model for classification and regression on tabular datasets that include text fields. Use its package to fit a pretrained model to your data, or its research tools to pretrain and evaluate on benchmarks.

cupy: Run NumPy and SciPy Workloads on GPUs
CuPy is a Python array library that brings NumPy- and SciPy-compatible computing to NVIDIA CUDA and AMD ROCm GPUs. It suits Python users who want GPU acceleration while reusing familiar APIs, or who need lower-level GPU controls.