PAPER DEEP DIVE
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
Tsinghua University and Beihang University present BSC-Nav (Brain-inspired Spatial Cognition for Navigation). The paper argues that existing embodied agents — whether end-to-end RL or "MLLM plus modular pipeline" — are fundamentally reactive and stateless: they process an observation and discard it, lacking any durable internal model of space, which yields fragmented knowledge, short-sighted planning and poor generalization. The authors borrow the answer from neuroscience, where spatial knowledge consolidates into three interconnected forms: landmarks, route knowledge, and survey knowledge. BSC-Nav instantiates these computationally in three modules. Landmark memory stores 4-tuples (world coordinates, open-vocabulary category, detection confidence, GPT-4o contextual description) with a spatial-overlap set plus confidence-weighted fusion for deduplication. The cognitive map extracts DINO-v2 patch features, projects them through inverse perspective projection and cascaded coordinate transforms into a voxel grid, and adopts a free-energy-principle-inspired surprise-driven update: a new feature is written when its mean distance to features in the n-hop neighborhood exceeds a threshold, replacing the lowest-surprise entry when the buffer is full, preserving cross-viewpoint diversity while bounding storage. Working memory retrieves hierarchically by instruction complexity — simple targets use text-only GPT-4 reasoning over landmark memory (even inferring unrecorded targets from co-located landmarks), while complex targets first have descriptions refined by GPT-4o, then "imagine" the appearance via Stable Diffusion 3.5, encoded by DINO-v2 and center-distance weighted pooled to query the cognitive map (imagine-then-localize), with similarity-weighted DBSCAN yielding candidate coordinates. Rather than greedily taking the highest confidence, candidates are ordered by H_i = lambda*p_i + (1-lambda)(1 - d_i/d_max). Across 62 MP3D/HM3D scenes and 8,195 episodes: OGN reaches 78.5% SR on HM3D (24.0 points above SOTA UniGoal), OVON zero-shot beats supervised DAgRL, IIN reaches 71.4%; SPL gains are even more consistent (IIN 57.2% vs 23.7%). On A-EQA it achieves the highest LLM-Match of 54.6, still trailing humans by 27.5. Real-world deployment on a custom platform (AgileX Ranger-mini-3.0 chassis, Franka Research 3 arm, RealSense D435i) ran 75 episodes in a ~200 m² two-floor space, with IIN reaching 100% SR on 4 of 5 targets and reliable localization to semantically plausible regions even on failure, plus long-horizon navigation-plus-manipulation demos such as "make breakfast" over three open-vocabulary objects. The paper proposes an embodied Turing test for spatial cognition probing three dimensions: real-time construction of reusable spatial representations, abstraction from sparse partial observations, and translation of high-level goals into actionable spatial plans.
Source: arXiv:2508.17198 (cs.AI) — Tsinghua University (Dept. of Computer Science and Technology; Dept. of Psychological and Cognitive Sciences) and Beihang University. Submitted Aug 24, 2025; 40 pages. Code: github.com/Heathcliff-saku/BSC-Nav.
In One Sentence
Whether they use end-to-end RL or "MLLM plus modular pipeline," today's embodied agents are fundamentally reactive and stateless — they process an observation and discard it, with no durable internal model of space, producing fragmented knowledge, short-sighted planning and poor generalization. BSC-Nav borrows the answer from neuroscience: biology consolidates spatial knowledge into three interconnected forms — landmarks, route knowledge, and survey knowledge. BSC-Nav instantiates that structure computationally with three modules: landmark memory encoding durable associations between salient cues and locations; a cognitive map that voxelizes egocentric trajectories into allocentric, map-like representations; and a working memory that hierarchically retrieves and composes both according to goal complexity. Coupled with MLLMs, it reaches state-of-the-art results across object-goal, open-vocabulary, text-instance and image-instance navigation with strong zero-shot generalization, and is validated on real-world navigation and mobile manipulation.
1. From Reactive to Cognitive
1.1 Where the Bottleneck Actually Is
Spatial cognition — acquiring, organizing, exploiting and updating knowledge about external space — is fundamental to both humans and AI. It underlies not only sensorimotor skills like navigation and manipulation but also higher functions including abstraction, planning and reasoning.
Despite rapid MLLM progress, current artificial systems remain fundamentally limited in large-scale spatial cognition, especially in long-horizon navigation and mobile manipulation. The central bottleneck is the lack of structured spatial memory — a mechanism for persistently encoding, organizing and retrieving spatial knowledge. Most existing methods "process observations in a reactive and stateless manner"; without durable internal models of external space, agents cannot consolidate coherent spatial representations or reason beyond immediate stimuli.
The authors therefore argue for a paradigm shift: from reactive processing to memory-centric spatial cognition, supporting persistent representations and compositional reasoning over time and space.
1.2 The Biological Three-Part Template
Decades of neuroscience show that organisms consolidate spatial knowledge into three distinct yet interconnected forms:
- Landmarks — encode stable associations of salient environmental cues to support localization and contextual understanding;
- Route knowledge — captures egocentric movement trajectories between landmarks for habitual navigation and path integration;
- Survey knowledge — integrates multiple routes into allocentric, map-like representations supporting flexible inference, shortcut discovery and detour planning.
These representations are accessed and coordinated via working memory — especially visual-spatial working memory — enabling adaptive retrieval, composition and generalization based on task demands and environmental familiarity.
RGB-D + pose"] --> L["Landmark Memory
(coords, category, confidence, description)"] O --> C["Cognitive Map
voxelized DINOv2 features
+ surprise-driven updates"] L --> W["Working Memory
hierarchical retrieval + exploration sequencing"] C --> W W --> P["Low-level planning
sim: Habitat shortest path
real: A* + TEB"] P --> V["Goal verification + affordance
360° scan + CLIP + GPT-4o"] V --> S["Higher-level skills
manipulation / EQA / instruction following"] end Bio -.->|inspires| Art
Figure 1: The BSC-Nav framework — (a) structured spatial memory in biological brains, (b) its computational instantiation, (c) navigation and higher-level spatial-aware skills it enables.
2. Method: Three Modules
2.0 Observation Space
Observations are $O_t = \{I_t, D_t, P_t\}$: RGB image $I_t \in \mathbb{R}^{H \times W \times 3}$, depth $D_t$, and pose $P_t = (X_t, Y_t, \phi_t^a) \in \mathbb{R}^3$ — where $X_t, Y_t$ is the 2D world position (defaulting to a Cartesian frame centered at the SLAM initialization point) and $\phi_t$ is the yaw angle. RGB feeds salient-object detection and visual feature extraction; depth and pose enable projection from pixel to world coordinates.
2.1 Landmark Memory
Landmark memory is a list of 4-tuples:
$$\mathcal{M}_{\text{landmark}} = \{L_k\}_{k=1}^{N},\quad L_k = \{c_k, \theta_k, \rho_k, T_k\}$$
where $\theta_k = (X_k^w, Y_k^w, Z_k^w) \in \mathbb{R}^3$ is the world coordinate of the $k$-th instance center, $c_k \in \mathcal{C}$ is an open-vocabulary category, $\rho_k \in [0,1]$ is detection confidence, and $T_k$ is a GPT-4o-generated description covering texture, shape and spatial context.
The pixel-to-world transform is spelled out completely. Given bounding-box center $(u_k, v_k)$ and depth $d_k$, inverse perspective projection gives the camera-frame point:
$$p_k^{\text{cam}} = d_k \cdot K^{-1}\begin{bmatrix} u_k \\ v_k \\ 1 \end{bmatrix} = d_k\begin{bmatrix} (u_k - c_x)/f_x \\ (v_k - c_y)/f_y \\ 1 \end{bmatrix}$$
with $K \in \mathbb{R}^{3\times 3}$ the intrinsic matrix. The pose then builds the base-to-world transform:
$$T^{\text{world}}_{\text{base}} = \begin{bmatrix} \cos\phi_t & -\sin\phi_t & 0 & X_t \\ \sin\phi_t & \cos\phi_t & 0 & Y_t \\ 0 & 0 & 1 & Z_{\text{base}} \\ 0 & 0 & 0 & 1 \end{bmatrix} \in SE(3)$$
and the point is mapped by cascaded homogeneous transformation $\begin{bmatrix} p_k^{\text{world}} \\ 1 \end{bmatrix} = T^{\text{world}}_{\text{base}} T^{\text{cam}}_{\text{base}} \begin{bmatrix} p_k^{\text{cam}} \\ 1 \end{bmatrix}$. In practice the open-vocabulary detector YOLO-World provides perception, with a predefined landmark category set $\mathcal{C} = \{\text{ sofa}, \text{ sink}, \text{ bed}, \dots\}$ and a confidence threshold excluding ambiguous or non-salient instances.
Deduplication by memory fusion is a pragmatically valuable design here. For a newly detected landmark $L_{N+1}$, define the spatial overlap set:
$$\mathcal{U} = \{L_j \in \mathcal{M}_{\text{landmark}} : \|\theta_{N+1} - \theta_j\|_2 < \delta_{\text{overlap}} \wedge c_{N+1} = c_j\}$$
If $\mathcal{U} \neq \varnothing$, fuse: $\theta_{\text{fused}}$ is the confidence-weighted average, $\rho_{\text{fused}}$ the mean confidence, and $T_{\text{fused}}$ the description of the highest-confidence member. Members of $\mathcal{U}$ are then removed and $L_{\text{fused}}$ added. Each 4-tuple therefore represents a unique spatial instance, avoiding information redundancy.
2.2 Cognitive Map
In parallel, the cognitive map continuously projects RGB observations along exploration trajectories into voxelized visual-spatial representations. It is a discrete voxelized structure:
$$\mathcal{M}_{\text{cog}} = \{F_v\}_{v \in \mathcal{V}},\quad F_v = \{f_b\}_{b=1}^{B},\ f_b \in \mathbb{R}^{\hbar}$$
where $\mathcal{V} \subseteq \mathbb{Z}^3$ is the voxel index space, each voxel $v = (v_x, v_y, v_z)$ maintains a buffer of up to $B$ feature vectors, and $\hbar$ is the feature dimension.
DINO-v2 extracts patch-level features $F_{\text{patch}} \in \mathbb{R}^{H' \times W' \times D}$ with $H' = H/s$, $W' = W/s$ and patch stride $s$. The center pixel of patch $(i,j)$ is:
$$(u_{ij}, v_{ij}) = (j \cdot s + s/2,\ i \cdot s + s/2)$$
Sampling depth $d_{ij}$ there and applying the same inverse projection and cascaded transform yields world coordinates, discretized into voxel indices:
$$v_{ij} = \left(\left\lfloor \frac{X^w_{ij}}{\Delta} + \frac{G}{2} \right\rfloor,\ \left\lfloor \frac{Y^w_{ij}}{\Delta} + \frac{G}{2} \right\rfloor,\ \left\lfloor \frac{Z^w_{ij}}{\Delta} \right\rfloor\right)$$
where $\Delta$ is voxel size, $G$ the grid dimension, and the $G/2$ offset centers the grid at the world origin.
2.3 Surprise-Driven Updates: Instantiating the Free-Energy Principle
The key question: should a new observation's feature be written into a voxel? Storing everything causes information overload and inefficient retrieval; conventional grid averaging or distance-weighted fusion introduces representational bias. BSC-Nav borrows the free-energy principle — biological systems refine internal models by minimizing prediction error. Computationally, this becomes a surprise score measuring deviation between new observations and existing memory:
$$S(f_{\text{new}}, v) = \frac{1}{|\mathcal{F}_{\mathcal{N}_n}(v)|} \sum_{f_b \in \mathcal{F}_{\mathcal{N}_n}(v)} D(f_{\text{new}}, f_b)$$
where $\mathcal{N}_n(v) = \{v' \in \mathbb{Z}^3 : \|v - v'\|_{\infty} \le n\}$ is the $n$-hop cubic neighborhood, $\mathcal{F}_{\mathcal{N}_n}(v)$ the union of feature buffers in that neighborhood, and $D(\cdot,\cdot)$ a distance metric such as cosine distance. With threshold $\tau$ (default 0.5): if $S > \tau$, add $f_{\text{new}}$ to $F_v$; if the buffer is full ($|F_v| = B$), replace the feature with the lowest surprise score.
Two benefits: first, selectively caching diverse features across viewpoints and timepoints enhances robustness in dynamic environments; second, avoiding redundant encoding of stable elements keeps storage and retrieval efficient. Maintaining a multi-feature buffer per voxel rather than a single aggregated vector is the core difference from conventional voxel maps.
Figure 2: Precise localization via hierarchical retrieval in working memory — (a) simple category-level goals query landmark memory, (b) complex instance-level goals use association-enhanced retrieval over the cognitive map.
2.4 Working Memory and Hierarchical Retrieval
Working memory activates only on receiving a navigation instruction, using a strategy guided by instruction complexity:
(a) Simple goals → MLLM-reasoning retrieval over landmark memory. Categories, confidences and contextual descriptions give MLLMs a basis for reasoning. Unlike direct rule matching, this enables context-aware inference that can even localize unrecorded targets — for instance, inferring where a "toaster" is from co-located landmarks like "stove" and "kitchen island" even if never explicitly recorded. Retrieval prompts guide a text-only GPT-4 to integrate confidence and descriptive semantics, producing candidate coordinates $\{\theta^{\text{cand}}_i\}_{i=1}^{K}$.
(b) Complex / fine-grained goals → association-enhanced retrieval over the cognitive map. To bridge the modality gap between text instructions and visual representations: text-only GPT-4o first refines the goal description with texture and spatial context, then a text-to-image model (default Stable Diffusion 3.5) "imagines" the target's likely appearance. This imagine-then-localize process resembles human pre-navigation thinking. The imagined image is encoded by DINO-v2 and pooled with center-distance weighting:
$$f_{\text{target}} = \frac{\sum_{i=1}^{N} w_i \cdot p_i}{\sum_{i=1}^{N} w_i},\quad \text{where } w_i = \exp\left(-\alpha \cdot \|(x_i, y_i) - (x_c, y_c)\|^2\right)$$
suppressing background interference and enhancing central target features. The pooled feature queries the cognitive map by cosine similarity, returning top-$K$ voxel coordinates, which are then passed through similarity-weighted DBSCAN whose cluster centers become the final spatial candidates.
(c) Exploration sequence planning. A crucial design: retrieval typically yields multiple candidates with different confidences and distances, and in category-level navigation where several valid targets exist, confidence alone is unreliable. Rather than greedily taking the highest confidence, BSC-Nav uses a composite score:
$$H_i = \lambda \cdot p_i + (1 - \lambda) \cdot \left(1 - \frac{d_i}{d_{\max}}\right)$$
where $p_i$ is existence probability (detection confidence for landmark memory, cosine similarity for the cognitive map), $d_i$ the Euclidean distance from the start, $d_{\max} = \max_j d_j$ for normalization, and $\lambda = 0.5$ by default — equal weight on existence likelihood and exploration efficiency.
2.5 Low-Level Planning, Goal Verification and Affordance
Low-level policies are environment-specific: in simulation, Habitat's greedy shortest-path algorithm operates on the scene's 3D mesh to compute optimal action sequences in a discrete action space. For real-world deployment a two-tier architecture is used — global planning with A\* over occupancy grids built by LiDAR SLAM, and local planning with the Timed Elastic Band (TEB) algorithm, which dynamically adjusts trajectories to avoid obstacles and outputs continuous velocity commands to the chassis.
On arrival: the robot performs a 360° rotational scan, computes cosine similarities between CLIP embeddings of the captured images and the goal's text or visual embedding, selects the best-aligned viewing angle, and feeds it to GPT-4o for precise verification. GPT-4o is also prompted to generate affordance-based actions guiding the robot to fine-tune pose, relative position and orientation — creating favorable initial conditions for downstream manipulation.
3. Experiments
Figure 3: Goal-directed multi-modal navigation — (a)(b) category-level and instance-level tasks, (c) trajectory visualizations with MLLM-assisted target verification.
3.1 Simulation: 62 Scenes, 8,195 Episodes
Evaluation runs in the Habitat simulator over physically reconstructed MP3D and HM3D environments, covering four navigation tasks. Agents start at random positions with discrete actions (forward 25 cm, turn left 30°, turn right 30°, stop); success requires executing stop within 1.0 m of the goal. Metrics are SR (success rate) and SPL (success weighted by path length).
| Task | Dataset | BSC-Nav SR | Best baseline | Gain |
|---|---|---|---|---|
| Object-Goal (OGN) | HM3D (6 categories) | 78.5% | UniGoal 53.1% | +25.4 |
| Object-Goal (OGN) | MP3D (20 categories) | 56.5% | UniGoal 52.6% | +3.9 |
| Open-Vocab (OVON) | MP3D unseen | 40.2% | DAgRL 37.1% (supervised) | zero-shot win |
| Open-Vocab (OVON) | MP3D seen | 38.9% | DAgRL 35.2% (supervised) | zero-shot win |
| Instance-Image (IIN) | HM3D | 71.4% | UniGoal 60.2% | +11.2 |
| VLN-CE R2R (long-horizon) | — | 38.5% | Uni-NaVid 37.4% | zero-shot |
| Task | BSC-Nav SPL | Best baseline SPL |
|---|---|---|
| OGN (HM3D) | 47.7% | UniGoal 30.4% |
| OGN (MP3D) | 30.2% | UniGoal 28.6% |
| IIN (HM3D) | 57.2% | UniGoal 23.7% |
| VLN-CE R2R | 53.1% | Uni-NaVid 35.9% |
Key points:
- 78.5% SR on HM3D OGN, 24.0 points above SOTA UniGoal; MP3D, with larger areas and 20 categories, still +15.5. The authors attribute this to UniGoal organizing only landmark memory with abstract goals and scene graphs, while BSC-Nav uses all three forms of spatial knowledge.
- On OVON with 79 everyday objects (e.g., "kitchen lower cabinet"), BSC-Nav zero-shot beats the supervised DAgRL.
- Efficiency gains are even more consistent than success gains — SPL is significantly higher across all tasks, thanks to the distance-plus-confidence scoring that lets the agent plan efficient exploration sequences by checking only a few candidate locations.
- On VLN-CE R2R long-horizon instruction following, 38.5% SR zero-shot comes within 8.5 points of the leading supervised method while being more efficient.
3.2 Active Embodied Question Answering (A-EQA)
On OpenEQA's A-EQA subset (184 questions across seven categories: Object Recognition, Object Localization, Attribute Recognition, Spatial Understanding, Object State Recognition, Functional Reasoning, World Knowledge), compared against a blind LLM, a question-agnostic frontier exploration strategy, and Explore-EQA. The metric is LLM-Match (semantic similarity between generated and reference answers computed by LLMs).
| Method | LLM-Match |
|---|---|
| Blind LLMs (GPT-4v / GPT-4o) | 35.5 / 35.9 |
| CG Cap. / SVM Cap. / Multi-Frame | 34.4 / 34.2 / 41.8 |
| Explore-EQA | 46.9 |
| BSC-Nav | 54.6 |
| Human | 82.1 (BSC-Nav trails by 27.5) |
By category, the strongest improvements are on the tasks requiring fine-grained spatial localization (OL and OSR) — consistent with the hypothesis that structured spatial memory enables rapid, accurate exploration. The authors are candid that a 27.5-point gap to humans remains, since human spatial cognition integrates commonsense knowledge, causal inference and abstraction from minimal environmental cues.
Figure 5: Universal navigation in real-world scenarios — (a) the custom mobile manipulation platform, (b) the ~200 m² two-floor environment, (c)-(h) performance across 15 targets and example trajectories.
3.3 Real World: 200 m² Two-Story Space plus Mobile Manipulation
A custom mobile platform: AgileX Ranger-mini-3.0 omnidirectional chassis + Franka Research 3 7-DoF arm + Intel RealSense D435i + LiDAR/IMU SLAM kit + RTX 4090 industrial computer, measuring 1528×773×494 mm.
In roughly 200 m² across two floors (office, lounge, reception, kitchen zones), 75 navigation episodes cover OGN / TIN / IIN with 5 targets each; each target tested over 5 trials from randomized starts, average path length 23.4 m:
- At least 3/5 successful trials per target (success = arriving within 1.0 m);
- IIN performs best, reaching 100% SR on 4 of 5 targets;
- OGN and TIN each reach 100% SR on 2 targets and exceed 66.7% in all cases;
- Even in failure cases BSC-Nav reliably localizes to semantically plausible regions — evidence of genuine spatial awareness rather than random wandering;
- Final Distance-to-Goal stays below 2.5 m in all cases with tight clustering near targets; mean velocity 0.76 m/s with low variance, indicating stable, efficient motion.
For mobile manipulation, natural language instructions are parsed by GPT-4 into waypoint-action sequences, with predefined manipulation primitives (grasp, place, pour, etc.) executed at each waypoint. Three demonstrations: single-waypoint (go to the marble table next to the shredder and clean the stain with a rag), object transfer (move a cookie box from the white table to the kitchen island, requiring navigation between manipulation actions), and multi-step ("make breakfast" — sequentially navigate to and manipulate three open-vocabulary targets: oatmeal jar, coco balls, milk bottle, assembling ingredients on a plate). The last one is the most telling: identifying and interacting with three spatially distributed open-vocabulary objects shows how critical structured spatial memory is for long-horizon reasoning and reliable object grounding.
4. Our Take
What is most admirable here is treating memory as a first-class design object rather than a cache. Landmarks, route knowledge and survey knowledge each have a clear biological counterpart, a clear data structure, and a clear update rule — that one-to-one cleanliness is rare in embodied AI papers.
Points with particular transfer value:
- Surprise-driven updating is a clever middle path. It avoids both extremes — store-everything and aggregation-fusion — by using mean neighborhood distance as a novelty criterion and keeping a multi-feature buffer instead of a single vector, preserving memory diversity while controlling storage. This applies to any long-term spatial memory system.
- Imagine-then-localize uses a text-to-image model for retrieval rather than generation — not to make pretty pictures, but to convert abstract text instructions into query vectors matchable against visual memory. A genuinely useful cross-modal retrieval trick.
- Not greedily taking the highest confidence but scoring $H_i = \lambda p_i + (1-\lambda)(1 - d_i/d_{\max})$ pays off directly in the large SPL leads. In multi-candidate settings, ranking by confidence alone makes agents backtrack repeatedly.
- The landmark fusion formula (confidence-weighted coordinates, mean confidence, description from the highest-confidence member) is a small but complete engineering detail addressing the duplicate-observation problem every SLAM or semantic mapping system hits.
Boundaries worth noting:
- The system is modular and zero-shot, so it leans on several off-the-shelf large models: YOLO-World for detection, DINO-v2 for features, GPT-4o for description and verification, SD 3.5 for imagination, CLIP for view selection. Inference cost and latency in real deployment are not negligible (per-step timing is not reported).
- A-EQA still trails humans by 27.5 points — an open challenge the authors acknowledge.
- The landmark category set $\mathcal{C}$ is predefined (sofa, sink, bed, etc.); open-vocabulary capability is carried mainly by YOLO-World and the cognitive-map branch, not fully unconstrained.
- Real-world evaluation is limited in scale: 75 episodes, 15 targets, a single 200 m² two-story environment, with manipulation using predefined primitives rather than a learned policy — what is demonstrated is integration of navigation with primitive invocation, not manipulation generalization.
Finally, the paper's proposal of an "embodied Turing test for spatial cognition" is worth remembering: future benchmarks should probe three foundational dimensions — (i) real-time construction of reusable spatial representations; (ii) abstraction and reasoning from sparse, partial observations; (iii) translation of high-level goals into actionable spatial plans. That characterizes "cognition" far better than SR and SPL alone.
5. Resources
- Paper: arXiv:2508.17198
- Code: https://github.com/Heathcliff-saku/BSC-Nav
- Affiliations: Tsinghua University (Dept. of CST; Dept. of Psychological and Cognitive Sciences; BNRist; THBI; Tsinghua-Bosch Joint ML Center) and Beihang University (Institute of Artificial Intelligence)
- Keywords: spatial navigation, spatial cognition, cognitive map, neuro-inspired learning, embodied intelligence
SOURCE LINKS