Open Source Computer Vision Projects
Discover 67 open source Computer Vision repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Computer Vision projects here are most often combined with Python, Machine Learning and Deep Learning. Last updated October 3, 2026.
67 repositories · updated October 3, 2026

pytorch-image-models: Use PyTorch Image Models and Training Tools
A Python library of PyTorch image models, pretrained weights, training utilities, and reference scripts. Use it to build image-classification systems, extract backbone features, or train and evaluate models across a broad range of architectures.

Kimi-k1.5: Train Multimodal LLMs with Reinforcement Learning
Kimi k1.5 is a research project describing reinforcement learning methods for training long-context, multimodal language models. It is aimed at researchers studying LLM reasoning and training, rather than users looking for a ready-to-install model.

Smart-AutoClicker: Automate Android Taps with Image Detection
Smart-AutoClicker, now presented as Klick'r, automates Android clicks and swipes using image detection or a simpler regular mode. It suits repetitive tasks and interaction testing where actions need to respond to on-screen elements.

CompreFace: Run Self-Hosted Face Recognition APIs
CompreFace is a Docker-based face analysis service with REST APIs for recognition, verification, detection, and related tasks. It suits teams that need to integrate face processing into applications while keeping deployment on their own infrastructure.
co-tracker: Track Points Across Video Frames
CoTracker tracks selected or grid-sampled points through video using a transformer-based model. It offers offline and online inference, pretrained checkpoints, and tools for evaluation and training.

Open-Interface: Control Your Computer with LLMs
Open-Interface turns natural-language requests into simulated mouse and keyboard actions, using screenshots and an LLM to guide and correct its work. It suits people who want to automate desktop tasks and are comfortable granting screen and input permissions.

notebooks: Learn and Apply Computer Vision Models
Roboflow notebooks is a hands-on tutorial collection for computer vision, covering model training, inference, detection, segmentation, and related tasks. Use it to explore techniques and run examples in hosted notebook environments.

TRELLIS: Generate 3D Assets from Text or Images
TRELLIS is a research model and toolkit for generating 3D assets from text or images. Its structured latent representation can produce meshes, 3D Gaussians, and radiance fields, but local use requires a compatible NVIDIA GPU and a substantial setup.

upscayl: Enlarge Images with AI
Upscayl is a desktop app for enlarging and enhancing low-resolution images with AI models. It runs on Linux, macOS, and Windows, and is best suited to users with a Vulkan-compatible GPU.

ImageToolbox: Edit, Convert, and Process Images on Android
ImageToolbox is a feature-rich Android app for photo editing, image conversion, OCR, PDF tasks, and other media workflows. It suits users who want many image utilities in one place, with distribution through Google Play, F-Droid, and GitHub releases.

wifi-3d-fusion: Sense Motion with WiFi Signals
WiFi-3D-Fusion captures WiFi CSI or RSSI data and visualizes motion in 3D, with optional research bridges for pose estimation and RF field reconstruction. It is aimed at researchers and experimenters with compatible hardware, not production or safety-critical use.

PaddleOCR: Extract Text and Structure from Images and PDFs
PaddleOCR is a Python toolkit for recognizing text and parsing document layouts in images and PDFs. It suits developers building OCR, document-processing, and retrieval workflows who need structured output and multilingual recognition.