smolvlm-realtime-webcam: Analyze Webcam Video with SmolVLM

Summary
A browser-based demo sends webcam frames to a llama.cpp server running SmolVLM 500M for real-time visual analysis. It suits developers exploring local vision models and webcam workflows, with performance depending on the model server and hardware.
At a glance
- Language
- HTML
- License
- NOASSERTION
- Stars
- 5.6k
- Forks
- 898
- Added to OSRepos
- October 11, 2025
- Last analyzed
- October 3, 2026
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
smolvlm-realtime-webcam is a small browser demo that connects a webcam to a SmolVLM model served by llama.cpp. It demonstrates a way to analyze camera input in real time, such as identifying objects, without building a larger application first.
It is most useful for developers testing multimodal inference and prompt-driven camera analysis. The repository provides a simple HTML entry point rather than a complete webcam product or model-serving stack.
Key Features
- Captures camera input through a browser page.
- Uses a llama.cpp server as the model-serving backend.
- Demonstrates SmolVLM 500M Instruct in GGUF format.
- Supports changing the instruction sent to the model, including requesting JSON output.
- Can be tried with other multimodal models supported by llama.cpp.
- Offers an optional GPU-offload setting for compatible hardware.
Use Cases
- Developers can quickly explore how a local vision-language model responds to webcam images.
- Prototypers can test prompts for object descriptions or structured visual output before building a larger interface.
- llama.cpp users can try a browser-based multimodal workflow against a running inference server.
- Educators and hobbyists can demonstrate camera-based AI analysis with a small model and a straightforward setup.
Project Facts
- Language: HTML
- License: NOASSERTION
- Stars: 5.6k
- Forks: 898
- Topics: none listed
- Archived: No
Getting Started
Install llama.cpp, then start the model server:
llama-server -hf ggml-org/SmolVLM-500M-Instruct-GGUF
Open index.html and use the page to start the camera demo. The README has the setup notes, including optional GPU offload and model alternatives.
Alternatives
- Qwen3-VL: Qwen3-VL supports image and video understanding with deployment examples, rather than a webcam demo built around SmolVLM and llama.cpp.
Considerations
- A llama.cpp server and the model must be available separately; this repository is the browser demo, not a bundled model runtime.
- The example is focused on experimentation and real-time visual analysis, not a production-ready application.
- GPU offload is optional, and the README notes the
-ngl 99setting for NVIDIA, AMD, or Intel GPUs. Actual performance can vary with hardware and model choice. - The repository metadata reports the license as
NOASSERTION, so verify licensing before reuse.
Found this useful?
Share it with someone who would like smolvlm-realtime-webcam.
Comparisons
Source repository
Open the original repository on GitHub.
35 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

markdown: Convert Markdown Text to HTML in Python
August 30, 2026
Python-Markdown converts Markdown text into HTML and supports extensions for additional syntax and behavior. It suits Python applications that need a configurable Markdown-to-HTML library rather than a standalone editor.

Awesome-Dynamic-Agent-Skills: Survey Evolving Agent Skills
August 22, 2026
A curated reading list and taxonomy for research on dynamic skill libraries in LLM agents. Use it to compare skill formats, lifecycle stages, update operators, verification approaches, and safety research.

jinja: Render Dynamic Templates in Python
August 5, 2026
Jinja is a fast, extensible template engine that turns templates and application data into documents. Python developers use it for HTML and other generated text, with features such as inheritance, autoescaping, async support, and sandboxing.

requests-html: Fetch and Parse HTML in Python
July 25, 2026
requests-html combines familiar Requests-style HTTP access with HTML parsing tools for scraping. It supports CSS and XPath selection, link extraction, asynchronous requests, and optional Chromium rendering for JavaScript-generated content.