OneCanvas: reproject video frames onto a single panoramic canvas so a stock 2D VLM reasons in 3D — SOTA on SQA3D, VSI-Bench and SPBench at an order of magnitude less compute

Matthias Niessner's team (@baranowskibrt, @davech2y, arXiv 2606.19253, 44-second demo video) presents OneCanvas: 3D scene understanding without complex geometry encoders or large training budgets — by compressing the scene onto a single image. The method: (1) extract per-patch features from all views with the frozen VLM vision encoder and unproject them to world coordinates using metric depth and camera poses; (2) place each patch on an equirectangular panoramic canvas at the continuous longitude/latitude of its 3D position as seen from the canvas origin — no rasterization onto a pixel grid, no cross-view aggregation, every patch its own token in one shared spatial coordinate system; (3) add a 3D position embedding to restore the depth lost when collapsing world position to angular coordinates. The pretrained VLM consumes the canvas like an ordinary image with no fusion or architectural changes. Since the canvas can be centered on any pose, the same representation directly supports situated reasoning from an agent's own viewpoint — a core robotics need. Bonus: a spatial pretraining curriculum that procedurally places real-image patch features at chosen 3D positions on an empty canvas — with appearance decoupled from size and location, geometry is the only solvable signal — generating unlimited annotation-free spatial supervision with controlled answer distributions to block shortcuts. Two-stage training (LoRA + position embedding for geometry, then merged with a fresh low-rank adapter for downstream formats). Results: SQA3D 65.3 EM@1 (+2.3 over prior SOTA), VSI-Bench 70.1 (+11.3 on route planning), SPBench zero-shot 72.1 (+4.8), all at an order of magnitude less training compute than the strongest competitors.





