Open Source Computer Vision Projects
Discover 67 open source Computer Vision repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Computer Vision projects here are most often combined with Python, Machine Learning and Deep Learning. Last updated October 3, 2026.
67 repositories · updated October 3, 2026

Waifu2x-Extension-GUI: Upscale Images and Video
A Windows desktop app for enlarging and denoising images, GIFs, and video, with AI-based video frame interpolation. It combines multiple processing engines and supports AMD, Nvidia, and Intel GPUs.

maestro: Fine-Tune Vision-Language Models
Maestro streamlines fine-tuning for multimodal vision-language models with ready-to-use recipes, a CLI, and a Python API. It suits developers adapting supported models to tasks such as JSON extraction and object detection.

GigaSLAM: Build Large-Scale Monocular Outdoor Maps
GigaSLAM is a monocular SLAM system for mapping and tracking in large, unbounded outdoor scenes using RGB video. It combines depth-assisted pose estimation, loop closure, and hierarchical Gaussian splats, and is aimed at research workflows with CUDA-capable GPUs.

MonoPCC: Estimate Monocular Depth in Endoscopic Images
MonoPCC is a PyTorch method for self-supervised monocular depth estimation on endoscopic images, using a photometric-invariant cycle constraint. It is aimed at researchers reproducing or extending depth and pose experiments on surgical-video datasets.

ai-baby-monitor: Local AI Video Monitoring for Baby Safety
A self-hosted baby-monitoring tool that checks webcam or RTSP video against natural-language safety rules using a local video LLM. It sends a quiet beep when a rule appears to be broken, with a dashboard for viewing the stream and model logs.

gaussian-splatting: Reconstruct and Render 3D Scenes in Real Time
The authors’ reference implementation of 3D Gaussian Splatting, which reconstructs scenes from posed images and renders novel viewpoints. It suits graphics and vision researchers and practitioners with CUDA-capable GPUs who need a trainable, interactive scene representation.

pymatting: Estimate Image Transparency for Clean Cutouts
PyMatting estimates foreground transparency from an image and a trimap, helping create clean cutouts for compositing. It offers several classical matting methods, foreground estimation, and a Python API and CLI.

PartCrafter: Generate Structured 3D Meshes from Images
PartCrafter generates part-separated 3D objects and scenes from a single RGB image using compositional latent diffusion. It suits researchers and developers exploring image-to-3D generation who have access to a CUDA-enabled GPU.

paper2gui: Run AI Models Through Desktop Apps
Paper2GUI is a desktop toolbox that packages AI models into ready-to-use apps for tasks such as image enhancement, speech synthesis, and video processing. It suits people who want to use these tools without building a development environment.

Lazyeat: Control Devices with Gestures and Voice
Lazyeat is a desktop tool for hands-free device control using camera-based gestures and voice input. It is aimed at people who want to operate a device while eating or whenever touching it is inconvenient.

rerun: Explore and Query Multimodal Robotics Data
Rerun is a data platform for logging, exploring, and querying synchronized visual and sensor data over time. It helps robotics and computer vision teams debug systems, inspect recordings, and prepare data for training.

big_vision: Train and Evaluate Large-Scale Vision Models
Google Research’s JAX and Flax codebase for training and evaluating vision and image-text models on GPUs and Cloud TPUs. It suits researchers running scalable experiments, but project-specific code may not stay compatible with the current core.