3D Reconstruction
3D reconstruction turns images, video, depth measurements, or other sensor data into a representation of an object or scene in three dimensions. It helps recover shape, scale, and spatial relationships when direct measurement is difficult, supporting tasks such as visualization, mapping, inspection, and simulation. Methods range from geometric techniques that match features across views to learned models that infer depth or generate new viewpoints. Their results may be represented as meshes, point clouds, volumes, or radiance fields.
Open source tools in this area include libraries for camera calibration, depth estimation, multi-view geometry, neural scene representations, and rendering. When choosing a tool, consider its license, maintenance activity, documentation, supported input data, hardware and runtime requirements, and compatibility with your existing workflow. Some approaches need many calibrated views or a powerful GPU, while others work from a single image. These tools are useful to researchers, developers, artists, and practitioners working with spatial data.
3 repositories · updated February 27, 2026

MonoPCC: Photometric-invariant Cycle Constraint for Monocular Depth Estimation
MonoPCC is a PyTorch implementation for monocular depth estimation, specifically designed for endoscopic images using a photometric-invariant cycle constraint. This self-supervised learning approach aims to improve depth prediction accuracy in challenging medical imaging scenarios. It demonstrates state-of-the-art performance on datasets like SCARED and KITTI, and offers a plug-and-play design for integration into various backbone networks.

gaussian-splatting: Reconstruct and Render 3D Scenes in Real Time
The authors’ reference implementation of 3D Gaussian Splatting, which reconstructs scenes from posed images and renders novel viewpoints. It suits graphics and vision researchers and practitioners with CUDA-capable GPUs who need a trainable, interactive scene representation.

VGGT: Visual Geometry Grounded Transformer for Rapid 3D Scene Reconstruction
VGGT, the recipient of the CVPR 2025 Best Paper Award, is a Visual Geometry Grounded Transformer developed by Facebook AI and the Visual Geometry Group at Oxford. This innovative feed-forward neural network efficiently infers key 3D scene attributes, including camera parameters, depth maps, and 3D point tracks, from single or multiple images within seconds. It offers a powerful solution for rapid 3D reconstruction and scene understanding.