Open Source Computer Vision Projects
Computer vision enables computers to interpret and analyze visual information, such as images and video. It supports tasks including object detection, image classification, facial analysis, text recognition, segmentation, and motion tracking. These techniques help automate inspection, improve accessibility, organize visual data, and give machines a way to respond to their surroundings. Methods range from traditional image processing to systems trained with machine learning and deep learning.
Open source tools in this area include libraries for image processing, model training and inference, data annotation, visual search, and real-time analysis. When choosing a tool, consider its license, maintenance activity, documentation, hardware and software requirements, and compatibility with your data and existing systems. These resources are useful to researchers, developers, educators, and organizations building or evaluating visual applications.
49 repositories · updated October 3, 2026

AREX-Skill: A Skill Library for Automated Machine Learning and Auto-Research
AREX-Skill is a powerful skill library designed to advance automated machine learning and auto-research. It distills over 5,000 executable skills from more than 1,000 popular GitHub repositories, making complex ML knowledge directly usable by coding agents. This project significantly enhances agent performance in various research tasks by providing structured, validated operating knowledge.

AWESOME-OCR-LLM: Curated Reading List for OCR in the LLM Era
AWESOME-OCR-LLM is a continuously updated reading list focusing on Optical Character Recognition (OCR) in the era of large language models (LLMs). It covers key areas like document parsing, understanding, visual text generation, and benchmarks, highlighting research from the past five years. This resource is invaluable for anyone tracking the rapid advancements in multimodal document AI.

scikit-video: Video Processing Routines for SciPy
scikit-video is a Python library designed for video processing, offering a suite of routines for tasks like I/O, quality metrics, and temporal filtering. Intended as a companion to scikit-image, it provides video-specific algorithms and aims for flexibility and GPU compute capabilities. This project offers a research-oriented alternative to existing frameworks, built entirely in Python.

Awesome-pytorch-list: Find PyTorch Libraries, Tutorials, and Papers
A categorized directory of PyTorch libraries, learning materials, and paper implementations. Use it to discover resources across NLP, computer vision, probabilistic modeling, and other deep-learning topics.

bg-remove: Remove Image Backgrounds in Your Browser
bg-remove is a browser-based image background remover that runs machine-learning models locally with Transformers.js. It suits people who want to edit images without uploading them to a server, with optional WebGPU acceleration where supported.

Qwen3-VL: Understand Images, Video, and Text with Multimodal Models
Qwen3-VL is a family of multimodal language models for interpreting images, video, and text. The repository provides inference examples, deployment guidance, and cookbooks for tasks such as OCR, spatial reasoning, and visual agents.

GLM-OCR: Recognize Text and Structure in Documents
GLM-OCR is a multimodal OCR model and SDK for extracting text and layout from complex documents. Use it through a hosted API or deploy the pipeline with supported inference servers for local control.

CellProfiler: Open-Source Software for Biological Image Analysis
CellProfiler is a powerful, free, and open-source application designed for biological image analysis. It empowers biologists without programming expertise to quantitatively measure phenotypes from thousands of images automatically. This tool simplifies complex image processing tasks, making advanced analysis accessible to a broader scientific community.

UI-TARS-desktop: Control Computers with a Multimodal AI Agent
UI-TARS-desktop is a desktop GUI agent that uses vision-language models to operate computers through natural-language instructions. It offers local and remote computer and browser operators for users building or exploring AI-driven automation.

PyTorch Image Models (timm): The Ultimate Collection of Image Encoders
PyTorch Image Models (timm) is an extensive library offering the largest collection of PyTorch image encoders and backbones. It provides a wide array of state-of-the-art models, complete with pretrained weights, training, evaluation, and inference scripts. This makes it an invaluable resource for researchers and developers working with computer vision tasks in PyTorch.

Smart-AutoClicker: Automate Android Taps with Image Detection
Smart-AutoClicker, now presented as Klick'r, automates Android clicks and swipes using image detection or a simpler regular mode. It suits repetitive tasks and interaction testing where actions need to respond to on-screen elements.

CompreFace: Run Self-Hosted Face Recognition APIs
CompreFace is a Docker-based face analysis service with REST APIs for recognition, verification, detection, and related tasks. It suits teams that need to integrate face processing into applications while keeping deployment on their own infrastructure.