Open Source Computer Vision Projects
Discover 67 open source Computer Vision repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Computer Vision projects here are most often combined with Python, Machine Learning and Deep Learning. Last updated October 3, 2026.
67 repositories · updated October 3, 2026

Isaac-GR00T: Adapt Vision-Language Models for Robot Control
NVIDIA Isaac GR00T is a vision-language-action model and toolkit for training and deploying generalist robot skills. It supports inference and fine-tuning across humanoid and other robot embodiments, with GPU hardware and model access requirements.
HunyuanVideo-Avatar: Create Audio-Driven Character Videos
HunyuanVideo-Avatar generates dynamic, emotion-controllable videos of one or more characters from avatar images and audio. It is aimed at creators and researchers who need expressive talking-avatar or dialogue video generation and have access to compatible NVIDIA GPU hardware.

OmniParser: Turn Screenshots into GUI Elements
OmniParser parses interface screenshots into structured elements, helping vision-based agents identify and ground actions on screen. It suits developers building computer-use agents that need visual UI understanding.

clarity-upscaler: Enhance and Upscale Images with AI
Clarity-Upscaler is a Python image-to-image project for increasing image resolution and enhancing details with Stable Diffusion workflows. It suits users comfortable with Cog or image-generation tools who want a configurable alternative to hosted upscaling services.

CineScale: Generate High-Resolution Video Without Fine-Tuning
CineScale is an inference framework for generating high-resolution video with pretrained diffusion models, without fine-tuning. It targets researchers and practitioners who want to upscale generation beyond a model’s training resolution, including 4K workflows.

Agent-S: Automate Desktop Tasks Through a GUI Agent
Agent-S is a Python framework that uses screenshots, mouse clicks, and keyboard input to carry out natural-language tasks in desktop applications. It suits research and automation workflows that need an agent to interact with a real computer interface.

StreamDiffusion: Generate Images in Real Time with Diffusion
StreamDiffusion adapts diffusion pipelines for interactive image generation, with support for text-to-image and image-to-image workflows. It targets developers building responsive GPU-powered demos and applications.

DragGAN: Edit GAN Images by Dragging Points
DragGAN is a research tool for interactive point-based editing of images generated by StyleGAN models. It lets users move image features by dragging points while the model updates the image, making controlled edits easier than conventional prompt-based workflows.

SyncTalk: Generate Synchronized Talking-Head Videos
SyncTalk is a CVPR 2024 system for generating talking-head videos from a person’s footage and audio. It targets synchronized lip movement, facial expression, and head pose, with workflows for training on a subject or running inference with provided models.

vggt: Reconstruct 3D Scenes from Images
VGGT is a feed-forward vision model that estimates cameras, depth, point maps, and point tracks from one or more images. It suits researchers and developers who need fast 3D scene reconstruction without a traditional multi-stage pipeline.

MuseTalk: Generate Audio-Synced Talking-Head Videos
MuseTalk creates lip-synced video from a source video or image and an audio clip using latent-space inpainting. It supports training and inference workflows, with real-time performance reported on a Tesla V100.

LAM: Create Animatable 3D Gaussian Avatars from One Image
LAM reconstructs a 3D Gaussian head avatar from a single image and supports animation and rendering across devices. It is aimed at developers building digital humans, especially interactive avatars, and requires model assets and a compatible compute environment for local use.