Open Source CUDA Projects
CUDA is a parallel computing platform and programming model for running general-purpose workloads on graphics processing units. It lets developers divide work across many GPU cores, speeding up tasks such as numerical computation, scientific simulation, image processing, and machine learning. CUDA tools can help move demanding workloads beyond the limits of CPU-only execution, while providing ways to manage memory, coordinate parallel tasks, and tune performance for specific hardware.
Open source tools in this area include GPU libraries, language bindings, compilers, kernel development utilities, and inference or computation frameworks. When choosing one, consider its license, maturity, maintenance activity, supported hardware and software versions, performance needs, and fit with existing code. These tools are useful to researchers, developers, and infrastructure teams building or optimizing GPU-accelerated applications.
4 repositories · updated September 26, 2026

KDA: Kernel Design Agents for High-Performance CUDA Kernel Development
Kernel Design Agents (KDA) offers an agent-centric workflow designed to streamline the research, implementation, verification, and iteration of performance-sensitive CUDA kernel tasks. This innovative approach leverages coding agents to accelerate the development of high-performance kernels. It is an early research prototype from NVlabs, welcoming community feedback and contributions.

ds4: A Fast Local Inference Engine for DeepSeek V4 Flash and PRO on Metal, CUDA, ROCm
ds4 is a highly optimized, native inference engine designed for DeepSeek V4 Flash and PRO models. It provides efficient local inference across various hardware platforms, including Apple Silicon (Metal), NVIDIA GPUs (CUDA), and AMD ROCm. This project focuses on delivering high performance for large language models on consumer-grade machines.

ZLUDA: Run CUDA Applications on Non-NVIDIA GPUs with Near-Native Performance
ZLUDA is an innovative open-source project providing a drop-in replacement for CUDA, enabling users to run CUDA applications on non-NVIDIA GPUs. Written in Rust, it aims to deliver near-native performance, significantly expanding hardware compatibility for CUDA-dependent software. This project offers a powerful solution for greater flexibility in GPU computing environments.

Numba: A Just-In-Time Compiler for Numerical Python Functions
Numba is an open-source, NumPy-aware optimizing compiler for Python, leveraging the LLVM project to generate machine code. It significantly accelerates numerical functions, offering support for automatic parallelization, GPU-accelerated code, and ufuncs. This tool is essential for Python developers seeking high-performance computing capabilities.