
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
Tsinghua University and Beihang University present BSC-Nav (Brain-inspired Spatial Cognition for Navigation). The paper argues that existing embodied agents — whether end-to-end RL or "MLLM plus modular pipeline" — are fundamentally reactive and stateless: they process an observation and discard it, lacking any durable internal model of space, which yields fragmented knowledge, short-sighted planning and poor generalization. The authors borrow the answer from neuroscience, where spatial knowledge consolidates into three interconnected forms: landmarks, route knowledge, and survey knowledge. BSC-Nav instantiates these computationally in three modules. Landmark memory stores 4-tuples (world coordinates, open-vocabulary category, detection confidence, GPT-4o contextual description) with a spatial-overlap set plus confidence-weighted fusion for deduplication. The cognitive map extracts DINO-v2 patch features, projects them through inverse perspective projection and cascaded coordinate transforms into a voxel grid, and adopts a free-energy-principle-inspired surprise-driven update: a new feature is written when its mean distance to features in the n-hop neighborhood exceeds a threshold, replacing the lowest-surprise entry when the buffer is full, preserving cross-viewpoint diversity while bounding storage. Working memory retrieves hierarchically by instruction complexity — simple targets use text-only GPT-4 reasoning over landmark memory (even inferring unrecorded targets from co-located landmarks), while complex targets first have descriptions refined by GPT-4o, then "imagine" the appearance via Stable Diffusion 3.5, encoded by DINO-v2 and center-distance weighted pooled to query the cognitive map (imagine-then-localize), with similarity-weighted DBSCAN yielding candidate coordinates. Rather than greedily taking the highest confidence, candidates are ordered by H_i = lambda*p_i + (1-lambda)(1 - d_i/d_max). Across 62 MP3D/HM3D scenes and 8,195 episodes: OGN reaches 78.5% SR on HM3D (24.0 points above SOTA UniGoal), OVON zero-shot beats supervised DAgRL, IIN reaches 71.4%; SPL gains are even more consistent (IIN 57.2% vs 23.7%). On A-EQA it achieves the highest LLM-Match of 54.6, still trailing humans by 27.5. Real-world deployment on a custom platform (AgileX Ranger-mini-3.0 chassis, Franka Research 3 arm, RealSense D435i) ran 75 episodes in a ~200 m² two-floor space, with IIN reaching 100% SR on 4 of 5 targets and reliable localization to semantically plausible regions even on failure, plus long-horizon navigation-plus-manipulation demos such as "make breakfast" over three open-vocabulary objects. The paper proposes an embodied Turing test for spatial cognition probing three dimensions: real-time construction of reusable spatial representations, abstraction from sparse partial observations, and translation of high-level goals into actionable spatial plans.