Open Source Computer Vision Projects

Computer vision enables computers to interpret and analyze visual information, such as images and video. It supports tasks including object detection, image classification, facial analysis, text recognition, segmentation, and motion tracking. These techniques help automate inspection, improve accessibility, organize visual data, and give machines a way to respond to their surroundings. Methods range from traditional image processing to systems trained with machine learning and deep learning.

Open source tools in this area include libraries for image processing, model training and inference, data annotation, visual search, and real-time analysis. When choosing a tool, consider its license, maintenance activity, documentation, hardware and software requirements, and compatibility with your data and existing systems. These resources are useful to researchers, developers, educators, and organizations building or evaluating visual applications.

49 repositories · updated October 3, 2026

AREX-Skill: A Skill Library for Automated Machine Learning and Auto-Research

AREX-Skill: A Skill Library for Automated Machine Learning and Auto-Research

AREX-Skill is a powerful skill library designed to advance automated machine learning and auto-research. It distills over 5,000 executable skills from more than 1,000 popular GitHub repositories, making complex ML knowledge directly usable by coding agents. This project significantly enhances agent performance in various research tasks by providing structured, validated operating knowledge.

PythonAutomated Machine LearningAI Agents
Added Sep 21, 2026 View details
AWESOME-OCR-LLM: Curated Reading List for OCR in the LLM Era

AWESOME-OCR-LLM: Curated Reading List for OCR in the LLM Era

AWESOME-OCR-LLM is a continuously updated reading list focusing on Optical Character Recognition (OCR) in the era of large language models (LLMs). It covers key areas like document parsing, understanding, visual text generation, and benchmarks, highlighting research from the past five years. This resource is invaluable for anyone tracking the rapid advancements in multimodal document AI.

OCRLLMDocument AI
Added Aug 11, 2026 View details
scikit-video: Video Processing Routines for SciPy

scikit-video: Video Processing Routines for SciPy

scikit-video is a Python library designed for video processing, offering a suite of routines for tasks like I/O, quality metrics, and temporal filtering. Intended as a companion to scikit-image, it provides video-specific algorithms and aims for flexibility and GPU compute capabilities. This project offers a research-oriented alternative to existing frameworks, built entirely in Python.

PythonVideo ProcessingScipy
Added Jul 27, 2026 View details
Awesome-pytorch-list: Find PyTorch Libraries, Tutorials, and Papers

Awesome-pytorch-list: Find PyTorch Libraries, Tutorials, and Papers

A categorized directory of PyTorch libraries, learning materials, and paper implementations. Use it to discover resources across NLP, computer vision, probabilistic modeling, and other deep-learning topics.

Awesome ListPyTorchMachine Learning
Added Jul 20, 2026 View details
bg-remove: Remove Image Backgrounds in Your Browser

bg-remove: Remove Image Backgrounds in Your Browser

bg-remove is a browser-based image background remover that runs machine-learning models locally with Transformers.js. It suits people who want to edit images without uploading them to a server, with optional WebGPU acceleration where supported.

TypeScriptReactMachine Learning
Added Jun 17, 2026 View details
Qwen3-VL: Understand Images, Video, and Text with Multimodal Models

Qwen3-VL: Understand Images, Video, and Text with Multimodal Models

Qwen3-VL is a family of multimodal language models for interpreting images, video, and text. The repository provides inference examples, deployment guidance, and cookbooks for tasks such as OCR, spatial reasoning, and visual agents.

AILLMMachine Learning
Added Jun 15, 2026 View details
GLM-OCR: Recognize Text and Structure in Documents

GLM-OCR: Recognize Text and Structure in Documents

GLM-OCR is a multimodal OCR model and SDK for extracting text and layout from complex documents. Use it through a hosted API or deploy the pipeline with supported inference servers for local control.

PythonOCRComputer Vision
Added May 28, 2026 View details
CellProfiler: Open-Source Software for Biological Image Analysis

CellProfiler: Open-Source Software for Biological Image Analysis

CellProfiler is a powerful, free, and open-source application designed for biological image analysis. It empowers biologists without programming expertise to quantitatively measure phenotypes from thousands of images automatically. This tool simplifies complex image processing tasks, making advanced analysis accessible to a broader scientific community.

PythonBiological Image AnalysisScientific Software
Added May 16, 2026 View details
UI-TARS-desktop: Control Computers with a Multimodal AI Agent

UI-TARS-desktop: Control Computers with a Multimodal AI Agent

UI-TARS-desktop is a desktop GUI agent that uses vision-language models to operate computers through natural-language instructions. It offers local and remote computer and browser operators for users building or exploring AI-driven automation.

TypeScriptAI AgentsDesktop App
Added May 6, 2026 View details
PyTorch Image Models (timm): The Ultimate Collection of Image Encoders

PyTorch Image Models (timm): The Ultimate Collection of Image Encoders

PyTorch Image Models (timm) is an extensive library offering the largest collection of PyTorch image encoders and backbones. It provides a wide array of state-of-the-art models, complete with pretrained weights, training, evaluation, and inference scripts. This makes it an invaluable resource for researchers and developers working with computer vision tasks in PyTorch.

PyTorchComputer VisionImage Classification
Added May 5, 2026 View details
Smart-AutoClicker: Automate Android Taps with Image Detection

Smart-AutoClicker: Automate Android Taps with Image Detection

Smart-AutoClicker, now presented as Klick'r, automates Android clicks and swipes using image detection or a simpler regular mode. It suits repetitive tasks and interaction testing where actions need to respond to on-screen elements.

AndroidAutomationComputer Vision
Added Apr 12, 2026 View details
CompreFace: Run Self-Hosted Face Recognition APIs

CompreFace: Run Self-Hosted Face Recognition APIs

CompreFace is a Docker-based face analysis service with REST APIs for recognition, verification, detection, and related tasks. It suits teams that need to integrate face processing into applications while keeping deployment on their own infrastructure.

Computer VisionJavaDocker
Added Apr 12, 2026 View details
Previous Page 1 Next

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️