VidMap at ECCV 2026: Video-Based Structure-from-Motion

Loading video
Loading videoVidMap (Pataki, Sarlin, Pollefeys; UZH/ETH CVG; ECCV 2026 poster #157, session #3) combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling metric reconstruction of arbitrary long uncalibrated videos. It leans on wide-baseline dense matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. On datasets with extreme motion and visual symmetries it is significantly more robust and accurate than state-of-the-art SLAM and SfM, classical or learned, with known or unknown calibration. The author also speaks at the ViLMa workshop on VidMap and follow-up work; code is public.
Structure-from-Motion3D reconstructionCamera Pose EstimationLoop ClosureMonocular DepthECCV2026SLAMComputer VisionVisual Localization
Category: perception
Author: @pesarlin
Date: 2026-09-07T00:00:00
Reference: https://github.com/cvg/vidmap





