Open Source Computer Vision Projects
Discover 67 open source Computer Vision repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Computer Vision projects here are most often combined with Python, Machine Learning and Deep Learning. Last updated October 3, 2026.
67 repositories · updated October 3, 2026

PPS-Ctrl: Translate Colonoscopy Images for Depth Estimation
PPS-Ctrl explores controllable sim-to-real translation for colonoscopy images, using per-pixel shading maps to guide Stable Diffusion and ControlNet. The repository currently provides partial pseudocode rather than a complete, ready-to-run implementation.

paperlib: Manage and Organize Academic Papers
Paperlib is a cross-platform desktop tool for collecting, searching, and organizing academic papers. It helps researchers find and correct publication metadata, manage PDFs and notes, and export references while writing.

deepface: Analyze and Recognize Faces in Python
DeepFace is a Python library that combines face recognition models and detection tools behind a simple API. Use it to verify identities, search face collections, generate embeddings, or estimate facial attributes.

audio2photoreal: Generate Photorealistic Avatars from Audio
A PyTorch research project for generating photorealistic human face and body motion from conversational audio. It includes person-specific pretrained models, training code, and access to the annotated dataset.

mlx-examples: Explore Machine Learning Models with MLX
A collection of standalone Python examples for the MLX machine learning framework, spanning language, image, video, audio, and multimodal models. Use it to learn MLX or adapt example implementations for experiments on supported hardware.

InfiniteTalk: Generate Audio-Driven Talking Videos
InfiniteTalk generates talking videos from audio and an image or existing video, synchronizing speech with facial expressions and body movement. It is aimed at creators and developers who need dubbed or audio-driven video, including long-form generation.

FlashVideo: Generate and Upscale High-Resolution Videos
FlashVideo is a two-stage text-to-video project that generates 270p clips and enhances them to 1080p. It is aimed at researchers and developers exploring efficient high-resolution video generation, with inference code and model weights available.

HunyuanWorld-1.0: Generate Interactive 3D Worlds from Text or Images
HunyuanWorld-1.0 turns text prompts or images into immersive, explorable 3D scenes. It combines panoramic image generation with semantic layering and 3D reconstruction, targeting creators and researchers who need editable world assets.

lance: Store and Query Multimodal Lakehouse Data
Lance is a Rust-based lakehouse format and SDK for AI and machine-learning data. It combines columnar storage with random access, vector and full-text search, and dataset versioning for teams working with multimodal data.

Step-Video-T2V: Generate Videos from Text Prompts
Step-Video-T2V is a 30-billion-parameter text-to-video model that generates clips up to 204 frames from English or Chinese prompts. It offers downloadable weights and inference code, but practical use requires substantial GPU memory and multi-GPU setup.

Wan2.2: Generate Videos from Text, Images, and Audio
Wan2.2 is a family of open video-generation models for text, image, and audio-driven workflows, plus character animation and replacement. It suits researchers and creators who can run large models on GPU hardware and want local inference options.

photopea: Edit Raster and Vector Graphics Online
Photopea is a browser-based editor for raster and vector graphics, including files from formats such as PSD, AI, and Sketch. Use it to edit images without installing desktop software, while keeping in mind that the editor itself is not fully open source.