Open Source Computer Vision Projects

Discover 67 open source Computer Vision repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Computer Vision projects here are most often combined with Python, Machine Learning and Deep Learning. Last updated October 3, 2026.

67 repositories · updated October 3, 2026

Waifu2x-Extension-GUI: Upscale Images and Video

Waifu2x-Extension-GUI: Upscale Images and Video

A Windows desktop app for enlarging and denoising images, GIFs, and video, with AI-based video frame interpolation. It combines multiple processing engines and supports AMD, Nvidia, and Intel GPUs.

C++Computer VisionMachine Learning
Added Mar 7, 2026 View details
maestro: Fine-Tune Vision-Language Models

maestro: Fine-Tune Vision-Language Models

Maestro streamlines fine-tuning for multimodal vision-language models with ready-to-use recipes, a CLI, and a Python API. It suits developers adapting supported models to tasks such as JSON extraction and object detection.

PythonMachine LearningComputer Vision
Added Mar 2, 2026 View details
GigaSLAM: Build Large-Scale Monocular Outdoor Maps

GigaSLAM: Build Large-Scale Monocular Outdoor Maps

GigaSLAM is a monocular SLAM system for mapping and tracking in large, unbounded outdoor scenes using RGB video. It combines depth-assisted pose estimation, loop closure, and hierarchical Gaussian splats, and is aimed at research workflows with CUDA-capable GPUs.

Computer VisionSlamC++
Added Feb 28, 2026 View details
MonoPCC: Estimate Monocular Depth in Endoscopic Images

MonoPCC: Estimate Monocular Depth in Endoscopic Images

MonoPCC is a PyTorch method for self-supervised monocular depth estimation on endoscopic images, using a photometric-invariant cycle constraint. It is aimed at researchers reproducing or extending depth and pose experiments on surgical-video datasets.

PythonMachine LearningDeep Learning
Added Feb 27, 2026 View details
ai-baby-monitor: Local AI Video Monitoring for Baby Safety

ai-baby-monitor: Local AI Video Monitoring for Baby Safety

A self-hosted baby-monitoring tool that checks webcam or RTSP video against natural-language safety rules using a local video LLM. It sends a quiet beep when a rule appears to be broken, with a dashboard for viewing the stream and model logs.

PythonAILLM
Added Feb 22, 2026 View details
gaussian-splatting: Reconstruct and Render 3D Scenes in Real Time

gaussian-splatting: Reconstruct and Render 3D Scenes in Real Time

The authors’ reference implementation of 3D Gaussian Splatting, which reconstructs scenes from posed images and renders novel viewpoints. It suits graphics and vision researchers and practitioners with CUDA-capable GPUs who need a trainable, interactive scene representation.

PythonComputer VisionComputer Graphics
Added Feb 13, 2026 View details
pymatting: Estimate Image Transparency for Clean Cutouts

pymatting: Estimate Image Transparency for Clean Cutouts

PyMatting estimates foreground transparency from an image and a trimap, helping create clean cutouts for compositing. It offers several classical matting methods, foreground estimation, and a Python API and CLI.

PythonLibraryComputer Vision
Added Feb 5, 2026 View details
PartCrafter: Generate Structured 3D Meshes from Images

PartCrafter: Generate Structured 3D Meshes from Images

PartCrafter generates part-separated 3D objects and scenes from a single RGB image using compositional latent diffusion. It suits researchers and developers exploring image-to-3D generation who have access to a CUDA-enabled GPU.

PythonGenerative AIMachine Learning
Added Jan 14, 2026 View details
paper2gui: Run AI Models Through Desktop Apps

paper2gui: Run AI Models Through Desktop Apps

Paper2GUI is a desktop toolbox that packages AI models into ready-to-use apps for tasks such as image enhancement, speech synthesis, and video processing. It suits people who want to use these tools without building a development environment.

AIMachine LearningDesktop App
Added Jan 7, 2026 View details
Lazyeat: Control Devices with Gestures and Voice

Lazyeat: Control Devices with Gestures and Voice

Lazyeat is a desktop tool for hands-free device control using camera-based gestures and voice input. It is aimed at people who want to operate a device while eating or whenever touching it is inconvenient.

Computer VisionDesktop AppCross Platform
Added Jan 6, 2026 View details
rerun: Explore and Query Multimodal Robotics Data

rerun: Explore and Query Multimodal Robotics Data

Rerun is a data platform for logging, exploring, and querying synchronized visual and sensor data over time. It helps robotics and computer vision teams debug systems, inspect recordings, and prepare data for training.

RustPythonComputer Vision
Added Jan 2, 2026 View details
big_vision: Train and Evaluate Large-Scale Vision Models

big_vision: Train and Evaluate Large-Scale Vision Models

Google Research’s JAX and Flax codebase for training and evaluating vision and image-text models on GPUs and Cloud TPUs. It suits researchers running scalable experiments, but project-specific code may not stay compatible with the current core.

Machine LearningDeep LearningComputer Vision
Added Dec 31, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️