1:12Astra Self-Calibrates Three Uncalibrated Cameras by Nudging the Arm@dhvanil · 100 views · 2026-09-08Camera CalibrationRobotic ArmSpatial understanding
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial UnderstandingMultimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, different models provide complementary spatial priors that benefit different tasks. Motivated by this, we propose ViPS, a novel multi-model prior framework designed to fully unleash the potential of incorporating multiple Visual Priors from diverse models into MLLMs for Spatial understanding. Specifically, ViPS introduces an Efficient Prior Proxy to generate multiple foundational priors with minimal inference overhead, and a Dynamic Prior Fusion mechanism to achieve harmonious and context-aware prior fusion and injection from the prior proxies. Extensive experiments demonstrate that ViPS successfully harmonizes diverse visual priors, establishing new state-of-the-art performance across multiple complex spatial reasoning and 3D spatial understanding benchmarks. Project page: https://visual-ai.github.io/vipsXiao Lin, Xiaohu Huang, Kai Han·Jul 16, 2026MLLMSpatial understandingVisual PriorJul 16, 2026