PAPER DEEP DIVE
LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding
DGIST team, accepted to IROS 2026. Robots operating long-term revisit evolving environments; existing systems either overwrite history to keep an up-to-date map or store semantic snapshots without cross-session object identity — the authors call this temporal amnesia: unable to answer "Where has the green chair been across all sessions?" LT-Mem is a three-layer solution: (1) Perception: MASt3R-SLAM multi-session alignment (Sim(3) anchors + inter-session loop closures + joint g2o optimization) plus SAM3 instance segmentation extracting per-object centroid/volume/visual embedding; (2) Reasoning: five evidence scores (E1 spatial proximity, E2 temporal continuity, E3 feature similarity as the primary signal w=0.45, E4 motion consistency, E5 occlusion handling) drive deterministic cross-session re-identification, with a constrained LLM judge only for ambiguous cases (MATCH/NEW-TRACK/HOLD); a structural integrity check triggers a session-wide HOLD on alignment failure; a volatility score V accumulates evidence in a Bayesian manner (Eq.2, one-time LLM prior), selecting among OVERWRITE/HOLD/MULTI-HYPOTHESIS updates; (3) Tri-Memory: Live stores current states, Delta logs timestamped events (APPEAR/DISAPPEAR/MOVE/RE-APPEAR/NONE), Meta accumulates volatility statistics feeding back into the update policy. The paper releases LT-VQA (two indoor envs, 10 objects × 10 sessions each, plus an outdoor parking lot over 10 sessions; 61 state-change events and 80 QA pairs). Results: Event F1 0.910 / QA-Event 0.820 / QA-Freq 0.600, beating Geometric/Text-Batch/VLM-Batch/STAR baselines across the board, with 438K tokens — 1/16 of VLM-Batch (7,114K) and 1/100 of STAR (45,859K); swapping in Qwen2.5-3B still reaches 0.885, showing the gains come from the structured memory architecture rather than LLM capacity. Ablations: removing re-identification collapses F1 to 0.140; removing volatility-awareness drops it to 0.615. Scene-level statistics on the parking lot even answer "when should I go to find a parking spot" (8:00 AM, 0 cars vs 1:03 PM, 34 cars).
TL;DR
Robots that operate for months in the same environment face objects that get moved, taken away, and put back. Existing systems either overwrite old states with the latest observation (the map is forever "now", history evaporates) or store semantic snapshots that cannot tell whether two sightings are the same chair. The DGIST team names this temporal amnesia — you can never answer "Where has the green chair been across all sessions?" LT-Mem (IROS 2026) solves it with a volatility-aware memory evolution framework: a perception layer provides spatially aligned object observations, a reasoning layer decides when and how memory should evolve, and a Tri-Memory structure keeps current states and full history side by side.
Sim(3) anchors + loop closures + g2o"] B --> C["SAM3 instance segmentation + lifting
centroid / volume / embedding"] C --> D{"Reasoning layer"} D --> E["Cross-session re-identification
5 evidence scores E1-E5
constrained LLM for ambiguity"] D --> F["Structural integrity check
anchor shift > 0.3m → session HOLD"] D --> G["Change detection
APPEAR / DISAPPEAR / MOVE
RE-APPEAR / NONE"] E --> H["Volatility V via Bayesian accumulation
(Eq.2,one-time LLM prior)"] G --> H H --> I["Update policy
OVERWRITE / HOLD / MULTI-HYP"] I --> J["Live Memory
current states"] I --> K["Delta Memory
timestamped event log"] I --> L["Meta Memory
volatility stats → feedback"] J & K & L --> M["Object-centric retrieval
→ long-horizon temporal QA"]

The problem: robots need a good memory too
Robots repeatedly revisit evolving environments: objects move, disappear, or are inconsistently observed due to occlusion and alignment noise. Multi-session SLAM provides a shared spatial frame, but spatial consistency does not guarantee valid object memory — session-indexed queries like "At Session 3 the robot dog was near the white desk; where is it at Session 9?" require persistent cross-session identity and explicit state-transition modeling, which no current system is designed for.
The paper builds a complete taxonomy of object-level events (Table I): APPEAR, DISAPPEAR, MOVE, NONE, RE-APPEAR, plus compound queries — trajectory, volatility, temporal localization, counterfactual. Each query type maps to a specific memory component; collapsing any component into a single map makes the corresponding query structurally unanswerable.
The dilemma of two research lines
Geometry-centric lifelong mapping (LT-mapper, Khronos) maintains global consistency under change but reduces object transitions to deletion and re-creation, losing temporal history. When an object relocates from A to B, the prior state is overwritten rather than recorded as a trajectory — the mechanism of temporal amnesia.
Semantic memory approaches (ReMEmbR's view-level embeddings, ConceptGraphs' open-vocabulary 3D scene graphs, 3D-Mem) build rich representations but rely on the latest observation or treat sessions independently, without enforcing persistent cross-session identity: two observations of a "chair" may or may not be the same physical instance. Reprocessing accumulated historical video streams also does not scale with prolonged deployment.
LT-Mem bridges the two: DUSt3R/MASt3R give metrically accurate multi-view geometry without task-specific training, SAM3 provides instance-level extraction across environments — yet the update dilemma (overwrite vs. accumulate) remains. The answer is coupling spatially grounded instances with a volatility-conditioned reasoning layer.
Perception layer: multi-session alignment + instance lifting
LT-Mem extends MASt3R-SLAM for multi-session alignment. Each session $S_t$ keeps keyframe poses $X_i^t \in \mathrm{Sim}(3)$ in a session-local frame; an anchor node $A^t \in \mathrm{Sim}(3)$ maps local to global, so the world-frame pose is $T_i^t = A^t \cdot X_i^t$. Scale comes from MR.ScaleMaster-style Sim(3) anchor estimation. Inter-session loop closures match keyframe images across sessions with MASt3R and compute relative Sim(3) constraints via ray-based geometric optimization; the factor graph (intra-session odometry edges + inter-session loop-closure edges) is jointly optimized with g2o.
Then, for each keyframe, SAM3 with text prompts yields 2D instance masks; masks are projected into the global frame using globally consistent poses and dense point maps to compute per-instance centroid $c_t^q$, volume $v_t^q$, and visual embedding $f_t^q$. Fragmented detections are merged by spatial consistency and point-cloud clustering into one observation per object per session. Crucially, all object centroids across sessions live in the same global frame, so displacement and disappearance become geometrically detectable.
Reasoning layer: evidence, checks, volatility, policy
Cross-session re-identification: deterministic first, LLM for the rest
Geometric proximity alone is insufficient for identity. LT-Mem computes five normalized evidence scores: E1 spatial proximity (weak prior; avoids penalizing displacement), E2 temporal continuity (stabilizing cue), E3 feature similarity (primary signal, weight 0.45, robust to viewpoint change), E4 motion consistency (conditioned on volatility), E5 occlusion handling (suppresses false disappearance). Experimental weights: 0.05/0.20/0.45/0.15/0.15. Candidates are pre-filtered by embedding retrieval (BGE-small-en-v1.5 + ChromaDB, no LLM in retrieval), top-5 per query. Simple cases are resolved by hard rules; ambiguous cross-session cases go to a constrained LLM judge whose output space is strictly {MATCH, NEW-TRACK, HOLD}. The LLM never modifies evidence scores — it only disambiguates over pre-computed structured inputs.
Structural integrity check: don't let bad alignment poison memory
Before updating, structural anchors verify alignment quality: mean anchor displacement $\bar{d}_{anchor}$ over $\tau_{align}=0.3$ m triggers a SESSION HOLD — all updates for that session are skipped and an alignment-failure event is recorded. Even suppressed updates leave an auditable trail. In all experiments every session passed the gate, but the mechanism guarantees errors do not propagate in the worst case.
Volatility: Bayesian evidence accumulation
The core variable is the volatility score $V_t \in [0,1]$:
$$V_t = \frac{P(E_t|V_{t-1})\,V_{t-1}}{P(E_t|V_{t-1})\,V_{t-1} + P(E_t|1-V_{t-1})(1-V_{t-1})}$$
where $E_t$ is modeled as a fixed monotonic function of $V$ (e.g., $P(\mathrm{MOVE}|V)$ increases with $V$). This is a lightweight temporal evidence accumulator, not learned. The initial prior $V_0$ comes from a one-time LLM query conditioned on the semantic label.
Fig. 5's four curves show the behavior: a persistently volatile brown basket; a fire extinguisher whose overestimated prior is progressively corrected; a green chair with mid-range fluctuation and partial recovery; a printer converging to near zero. Even when the LLM prior is wrong, $V_t$ converges within a few sessions through the Bayesian update — no per-object tuning needed. This matters because volatility is not a class-level fixed property: the same object type behaves differently in different environments.
Deterministic update policy
Given the change decision and volatility, a deterministic policy selects: HOLD (preserve under low confidence or occlusion), OVERWRITE (commit the new state), or MULTI-HYPOTHESIS (keep competing location hypotheses with confidence scores; a hypothesis is promoted when it beats competitors by margin $\Delta$ over consecutive sessions). Change detection is not a single threshold but a priority-ordered deterministic rule set: alignment quality → missing-evidence signals (disappearance from the old location) → volatility-conditioned motion tolerance. Volatile objects naturally get a larger motion tolerance and are not misjudged as disappear-reappear pairs.
Tri-Memory: not an independent invention, a natural consequence
The authors are explicit: Tri-Memory is not proposed as an independent architectural contribution — it follows directly from the volatility-aware reasoning layer, which produces three heterogeneous outputs (current state estimate, timestamped event record, long-term statistics) that cannot be collapsed into one map representation without reintroducing temporal amnesia.
- Live Memory: the current confirmed state per object. Present-state queries ("Where is the white chair now?") are answered here.
- Delta Memory: a timestamped per-object event log — MOVE/APPEAR/DISAPPEAR/RE-APPEAR/NONE with session index and metric context. Alignment-failure events are logged too. History and counterfactual queries are answered by traversing this log.
- Meta Memory: long-term statistics (volatility $V_t$, change frequency) that feed back into the reasoning layer by conditioning future motion tolerance — closing the observation↔policy loop.
The authors also trace terminology: in lifelong mapping, delta/meta maps denote geometric point-cloud differences (LT-mapper). LT-Mem's Delta/Meta Memories operate at the object-semantic level: Delta records identity-preserving transitions ("this specific chair moved from A to B at S5"), Meta accumulates per-object behavioral statistics. This shift from point-level differencing to object-level event logging is what enables structured temporal queries. At query time, an object-centric retrieval module routes each question to the relevant components; metric queries (displacement, frequency) are answered by deterministic computation over the event log — numerically grounded, independent of LLM hallucination.

LT-VQA: a benchmark built for the long term
Existing datasets don't fit: 3RScan has instance correspondence across rescans but no event-level annotations or temporal QA; NaVQA is single-session. LT-VQA jointly provides globally aligned multi-session geometry, persistent object identity, event-level transition annotations, and session-indexed temporal QA. The authors note that cross-session event annotation requires meticulous per-object multi-temporal alignment — inherently resource-intensive, which is why such a dataset didn't exist.
| Environment | Task | Sessions | Frames/Session | Duration/Session |
|---|---|---|---|---|
| Lab-S (compact indoor, 10 objects) | Instance history | 10 | 116–159 | 1.6–2.1 min |
| Lab-L (larger indoor) | Instance history | 10 | 115–183 | 2.6–3.3 min |
| Parking Lot (outdoor) | Spatial statistics | 10 | 124–195 | 2.2–3.2 min |
Objects range from highly volatile (brown basket, scissors) to fully stationary (fridge, sofa). Annotation statistics: 61 state-change events and 80 QA pairs across the two indoor environments. Capture is monocular RGB (iPhone 15/17 Pro); all experiments ran on an RTX 5070 Ti with 64 GB RAM, LLM components unified to Gemini 2.5 Pro for fair comparison.

Experiments: leading across the board, two orders of magnitude cheaper
Metrics and baselines
Instance history tracking reports Event F1 (each ground-truth event as a positive instance), QA-Event (exact-match accuracy on event-history queries) and QA-Freq (frequency queries against Meta Memory statistics); QA outputs are structured JSON matched against annotations. Baselines deliberately isolate reasoning paradigms rather than retrofit existing systems (Khronos/ConceptGraphs/ReMEmbR produce no object-level event logs; augmenting them would go beyond their design scope):
- Geometric Only: distance-thresholding change detection; no event log, so history/frequency queries are structurally unanswerable (N/A).
- Text-Batch: per-keyframe captions aggregated into session summaries; no identity, no spatial grounding.
- VLM-Batch: all keyframes per object in a single batch query; full visual context, purely implicit VLM reasoning.
- STAR: a recent agentic memory-action retrieval system, run directly with embodied navigation actions disabled.
Main results
| Method | Event F1↑ | QA-Event↑ | QA-Freq↑ | Tokens↓ |
|---|---|---|---|---|
| Geometric Only | 0.630 | N/A | N/A | – |
| Text-Batch | 0.460 | 0.300 | 0.267 | 3,973K |
| VLM-Batch | 0.790 | 0.680 | 0.333 | 7,114K |
| STAR† | 0.420 | 0.460 | N/A | 45,859K |
| Ours (Qwen2.5-3B) | 0.885 | 0.800 | 0.567 | 352K |
| Ours (Gemini 2.5) | 0.910 | 0.820 | 0.600 | 438K |
†STAR emits free-form text, scored by keyword match; all others use strict exact-match. Points worth chewing on:
- STAR consumed the most tokens (45,859K) and scored the lowest Event F1 (0.420): without persistent identity and an event log, per-query retrieval over raw records cannot compose coherent object histories; frequency queries are unanswerable.
- VLM-Batch shows context length can't buy structure: the fullest visual context reaches 0.790/0.680, at 16× LT-Mem's token cost.
- Swapping in Qwen2.5-3B still reaches 0.885/0.800/0.567: the gains come from the structured memory architecture, not LLM capacity — the paper's most important conclusion, further supported by ablations.
- Per-session token budget is independent of frame count: perception is a one-time offline cost; the reasoning layer sees only compact text. The gap widens as sessions grow.

Spatial statistics: memory's "dimensionality-reduction strike"
The Parking Lot sequence (10 sessions across different days and times) evaluates aggregate scene-level patterns. With per-session vehicle counts in Meta Memory, "When should I go to find a parking spot?" and "When is it hardest to park?" are answered by deterministic lookup: Session 3 (8:00 AM) had 0 cars — go then; Session 8 (1:03 PM) had 34 — avoid it. No additional inference.
Ablations: identity is the foundation, volatility is the precision
| Configuration | Event F1 | QA-Event | QA-Freq |
|---|---|---|---|
| w/o Re-ID (each observation an independent track) | 0.140 | 0.120 | 0.070 |
| w/o Volatility (fixed threshold) | 0.615 | 0.640 | 0.130 |
| Full | 0.910 | 0.820 | 0.600 |
Removing cross-session re-identification collapses everything — persistent identity is the foundational prerequisite for all downstream temporal reasoning. Replacing volatility-conditioned updates with a fixed threshold degrades all metrics: thresholds cannot adapt to per-object dynamics and false events corrupt Meta Memory statistics.

Limitations and outlook
The authors acknowledge three points: LT-VQA currently spans two indoor and one outdoor environment over 10 sessions each; cross-session re-identification remains hard when multiple objects share similar appearance and spatial context — identity preservation under ambiguity is a fundamental difficulty; scaling to longer deployment horizons and establishing LT-VQA as a public benchmark are future directions.
Why this paper is worth a careful read
- It reframes "long-term autonomy" from a mapping problem to a memory problem: spatial consistency is necessary but not sufficient; the core question is when and how object states should evolve, adaptively per object.
- The LLM is used with remarkable restraint: one prior, ambiguity adjudication, answer organization — all numeric conclusions come from deterministic computation, directly cashed out in the 1/100 token cost.
- The ablation design is persuasive: dual runs with Qwen2.5-3B and Gemini prove architectural gains; the w/o Re-ID collapse quantifies that "identity first" is the right priority order.
- For teams working on SLAM, scene graphs, or embodied memory, LT-VQA offers a previously missing longitudinal benchmark with event-level annotations.
Project page: lt-mem.github.io
References
- Paper: arXiv:2608.19059 — LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding (IROS 2026)
- Related systems: Khronos (RSS 2024), LT-mapper (ICRA 2022), ReMEmbR (ICRA 2025), ConceptGraphs (ICRA 2024), MASt3R-SLAM (CVPR 2025), SAM 3 (arXiv:2511.16719), STAR (ICRA 2026)
- Dataset: LT-VQA (lt-mem.github.io)
SOURCE LINKS



