Open Source Machine Learning Projects
Discover 191 open source Machine Learning repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Machine Learning projects here are most often combined with Python, Deep Learning and LLM. Last updated October 4, 2026.
191 repositories · updated October 4, 2026

PartCrafter: Generate Structured 3D Meshes from Images
PartCrafter generates part-separated 3D objects and scenes from a single RGB image using compositional latent diffusion. It suits researchers and developers exploring image-to-3D generation who have access to a CUDA-enabled GPU.

EliteQuant: Find Quantitative Finance Resources
EliteQuant is a curated directory of online resources for quantitative modeling, trading, and portfolio management. It helps practitioners and learners find tools, data sources, research, and communities across the quantitative finance ecosystem.

csm: Generate Conversational Speech from Text and Audio
CSM is Sesame’s speech-generation model, producing audio from text and optional conversation context. It suits developers building voice experiences who can run large models on a CUDA-compatible GPU and provide the required Hugging Face checkpoints.

paper2gui: Run AI Models Through Desktop Apps
Paper2GUI is a desktop toolbox that packages AI models into ready-to-use apps for tasks such as image enhancement, speech synthesis, and video processing. It suits people who want to use these tools without building a development environment.

modular: Build and Deploy AI Models with MAX and Mojo
Modular combines the MAX framework for AI development and deployment with Mojo, a programming language and compiler. It suits developers building model-serving systems, accelerator code, or software using Mojo.

optimum: Optimize Model Training and Inference on Target Hardware
Hugging Face Optimum adds tools for optimizing model training and inference across hardware backends. It suits teams using Transformers, Diffusers, TIMM, or Sentence Transformers who need hardware-specific deployment or training workflows.

litgpt: Train, Fine-Tune, and Deploy Large Language Models
LitGPT provides implementations and workflows for pretraining, fine-tuning, evaluating, and serving a range of large language models. It suits developers and researchers who want configurable training recipes and direct control over model code.

Qwen3-Coder: Generate Code and Power Coding Agents
Qwen3-Coder is Qwen’s family of code-focused language models for code generation, repository-scale understanding, and agentic development tasks. It offers open-weight checkpoints in several sizes, with context lengths up to 256K tokens.

TabSTAR: Apply a Tabular Foundation Model to Data with Text Fields
TabSTAR is a Python model for classification and regression on tabular datasets that include text fields. Use its package to fit a pretrained model to your data, or its research tools to pretrain and evaluate on benchmarks.

Paper2Code: Generate Code Repositories from ML Papers
Paper2Code is a research project that uses specialized LLM agents to turn machine-learning papers into code repositories. It suits researchers and developers exploring paper reproduction, with setup options for OpenAI APIs or vLLM.

transformerlab-app: Train and Evaluate AI Models in One Workspace
Transformer Lab is a research workspace for training, fine-tuning, running, and evaluating AI models on local machines or GPU clusters. It suits individual researchers who want a unified UI and teams that need to coordinate jobs across existing infrastructure.

big_vision: Train and Evaluate Large-Scale Vision Models
Google Research’s JAX and Flax codebase for training and evaluating vision and image-text models on GPUs and Cloud TPUs. It suits researchers running scalable experiments, but project-specific code may not stay compatible with the current core.