PAPER DEEP DIVE
Robot Planning and Situation Handling with Active Perception (VAP-TAMP)
SUNY Binghamton + CMU + Ford Research + Agility Robotics. A TAMP (task and motion planning) framework with VLM-based active perception. Core problem: TAMP assumes a fully observable world, but execution constantly surfaces unforeseen situations — a door only half open, one lemon half fallen out of the plate. Prior open-world planning implicitly assumed situations are fully observable, while real robots rely on unreliable onboard sensors. VAP-TAMP’s key novelty is the bidirectional interaction between VLMs and action knowledge: predicates from the current action’s PDDL preconditions/effects construct VQA prompts for the VLM, and VLM outputs update the action knowledge. Three modules: (1) scene graph generator (RGB-D point cloud + instance segmentation + CLIP embeddings; geometric rules for on/inside/near; updated via both action-driven and observation-driven channels); (2) active perception (N semantically equivalent paraphrases per predicate with majority voting; inconsistent responses → the VLM suggests a viewing direction (left/right/closer/above) → navigate and re-observe, looping within budget K until the view is sufficient); (3) situation handling (verify-execute-verify loop: preconditions checked before, effects after; any failure → correct the scene graph → PDDL replanning; diagnoses whether discrepancies stem from perception error, world state change, or execution failure). Real-world: Segway RMP + UR5e + Robotiq gripper, 4 household tasks × 25 trials (100 total), 88% success (OK-Robot 76%, COWP 69%, Closed-World 33%); 72/100 trials contained at least one unforeseen situation; VLM-guided viewpoint selection averages 1.61 views vs 4.35 for greedy exploration (63% fewer); failure-mode analysis shows VAP-TAMP failures concentrate in manipulation hardware (42%) — the perceptual bottleneck has been shifted away. Simulation (OmniGibson, 5 tasks with injected failure probabilities) shows predicate-based verification beats SuccessVQA/AffordanceVQA and their combination, with effect verification contributing more than precondition verification (grasp/cut carry 50% combined failure rates detectable only post-execution).
TL;DR
A robot's plan can be perfect and still fail because the world changed during execution: a door only half open, one lemon half fallen off the plate, a cup moved before the grasp. TAMP (task and motion planning) methods generally assume perfect state estimation and static environments; encountering an unforeseen situation, they halt or ask a human. This IROS work — VAP-TAMP from SUNY Binghamton with CMU, Ford Research, and Agility Robotics — offers a solution: use PDDL action knowledge (preconditions/effects) to construct prompts for a VLM, letting the robot actively select viewpoints, verify predicates, diagnose situations, and replan. 88% success over 100 real-world household task trials.
RGB-D point cloud + segmentation + CLIP
on / inside / near predicates"] SG --> P["PDDL planner
action sequence"] P --> EXE["Execute action a"] EXE --> PRE{"Precondition verification
VerifyPredicate"} PRE -- "False" --> CORR["Correct scene graph"] --> P PRE -- "True" --> ACT["Execute + apply effects"] ACT --> EFF{"Effect verification"} EFF -- "False" --> CORR EFF -- "True" --> NEXT["Next action"] PRE & EFF --> AP["Active perception
N paraphrases, majority vote
inconsistent → VLM suggests view
navigate & re-observe within budget K"]

The problem: three challenges of the open world
- Perception is unreliable: occlusion, clutter, and poor angles produce predicate errors — and VLMs provide no confidence signal.
- The world changes between planning and execution: objects get displaced by humans, paths get blocked, plans computed from initial observations become invalid.
- Actuation is imperfect: grasps slip, placements miss. Classical systems halt or require intervention.
The critique of prior open-world planning is pointed: knowledge representations for explaining failures, LLM-augmented action knowledge (COWP), asking humans when unsure — all of these implicitly assume situations are fully observable, jumping straight to "handling" and skipping "reliably perceiving" the situation.
The contrast with VLM-based work is equally key: SayCan / Code as Policies / PaLM-E / VoxPoser treat VLM outputs as one-shot, infallible oracles — query once, trust implicitly, never verify. VAP-TAMP does the opposite: it turns VLM uncertainty into an actionable signal that triggers deliberate information gathering. Classical next-best-view maximizes geometric coverage but cannot tell which viewpoint would resolve VLM uncertainty about semantic predicates like containment; interactive perception physically reveals occluded surfaces but alters the world. VAP-TAMP takes a third path: let the VLM say where to look.
Historically, active perception is an old theme in robotics (Bajcsy 1988), but its classical form addressed geometry — which views are needed to reconstruct an object. The LLM/VLM era changed the question: not "is the view complete enough" but "can the truth value of a symbolic predicate be reliably determined". VAP-TAMP moves active perception from geometric space to symbolic space, replacing geometric coverage with predicate-verification confidence as the optimization target.

Method: three modules and one closed loop
Scene graph generator: gatekeeper of symbolic state
The system builds an initial scene graph from RGB-D exploration: observations aggregate into a point-cloud map, instance segmentation partitions it into objects $O$, each stored with a 3D centroid, bounding box, and CLIP embedding for open-vocabulary retrieval. The relationship set $R$ is evaluated geometrically — on requires the centroid above a support surface within a threshold, inside requires bounding-box containment, near requires centroid distance below a threshold.
Maintenance runs on two channels: action-driven updates (executing pick(cup, table) immediately removes on(cup,table) and adds holding(robot,cup)) and observation-driven updates (when verification detects a mismatch with physical reality, correct the graph — which triggers replanning). The PDDL planner therefore always searches over an accurate world model.
Active perception: turning VLM hesitation into motion
The VerifyPredicate procedure: (1) generate $N$ semantically equivalent natural-language questions for the target predicate ("Is the cup on the table?", "Is the cup resting on the table surface?", "Is the cup positioned on top of the table?"); (2) query the VLM per question, take a majority vote; (3) if responses are consistent, ask the VLM whether the current view is sufficient — if yes, accept; (4) if inconsistent or insufficient, have the VLM suggest a direction (left/right/closer/above…) and navigate there, looping until sufficient or budget $K$ is exhausted.
The core insight: VLMs reason far more reliably about concrete observable state predicates ("Is the robot holding the cup?") than about abstract action outcomes ("Did the grasp succeed?") — quantitatively confirmed in simulation. And viewpoint selection is not geometric-coverage-driven next-best-view; the VLM itself says where it needs to look.
Situation handling: verify-execute-verify
Each action follows a three-phase pattern: verify all preconditions, execute, verify intended effects. Any failure → correct the scene graph from observation → PDDL replanning. A "situation" is defined as an unexpected world state that prevents task completion using a plan that would normally succeed; sources are environmental change, perceptual error, and execution failure. The framework's design goal is exactly to diagnose the root: perception errors are fixed by active perception, world-state changes require plan adaptation, execution failures require recovery actions.

How action knowledge "teaches" the VLM: the journey of one PDDL snippet
Section IV of the paper gives a complete example. The domain's pick action reads:
:action pick :parameters (?a - robot ?o - object ?s - surface) :precondition (and (on ?o ?s) (reachable ?o) (hand_empty ?a)) :effect (and (holding ?a ?o) (not (hand_empty ?a)) (not (on ?o ?s)))
Before execution, the situation handler decomposes this structure predicate by predicate into concrete visual questions: "Is the cup on the table?", "Is the robot's hand empty?" If hand_empty evaluates False (the hand still holds something else), the framework does not force the grasp: it first corrects the scene graph from observation (perhaps the cup is actually in the cabinet), then replans from the corrected state — possibly targeting a different location entirely. After execution it verifies holding(robot, cup): if the grasp slipped and the cup fell, effect verification catches it immediately, the graph records "cup on floor", and replanning may insert a pick(floor) recovery. No step needs human intervention.
Action table: the checks every action must pass
| Action | Preconditions | Situations to detect and recover from |
|---|---|---|
| find | Object and agent in the same room | Object not in view after navigation; no free space near the object; held object drops during navigation |
| grasp | Object in view; hand empty | Grasp fails, object unchanged; grasp fails, object drops nearby |
| placein | Object in hand; receptacle in view; receptacle open | Place fails, object remains in hand; place fails, object drops nearby |
| placeon | Object in hand; receptacle in view | Same two failure modes |
| open / close | Object in view | Open/close fails, object unchanged |
| turnon | Object in view | Turn-on fails, object remains off |
| cut | Object in view; knife in hand | Object not cut, knife in hand; object not cut, knife drops nearby |
Note that every "situation" is a concrete, observable physical state ("object drops nearby", not "the task went wrong") — which is exactly why predicate verification can detect them one by one.
The scene graph's dual-ledger mechanism
The two update channels deserve their own paragraph. Action-driven updates are an optimistic ledger: assume success, write the effects into symbolic state immediately, so the planner can proceed. Observation-driven updates are a pessimistic reconciliation: when effect verification finds holding(robot,cup) is actually False, correct the ledger from observation and report the correction as a situation triggering replanning. The tension between the two ledgers is the framework's sensitivity: optimistic-only accounting (as in many pure LLM-VLA systems) lets failures accumulate silently until some action becomes impossible; pessimistic-only accounting requires full re-observation after every action and collapses in efficiency. Verify-execute-verify bounds reconciliation cost to each action's effect predicates, balancing timeliness and efficiency.
The budget philosophy of active perception
The loop bound $K$ (viewpoint budget) is an easily overlooked but engineering-critical parameter. A real robot cannot "take one more look" indefinitely — every move costs time and energy, and some predicates are physically unresolvable (a cup locked in a closed cabinet has no visible inside state from any view). When the budget is exhausted the algorithm returns the current majority vote, and the system proceeds with "the best judgment after K observations": either execute and rely on effect verification, or recover via situation handling after failure. This contrasts with "ask a human when the LLM is unsure" (Ren et al.): VAP-TAMP first tries active perception to disambiguate, leaving unresolvable uncertainty to the closed-loop recovery mechanism rather than interrupting the task flow.
Real-world experiments: 100 trials, 88% success
Platform and tasks
Segway RMP base + UR5e arm + Robotiq 2F-85 gripper + wrist-mounted RealSense D435i; VLM queries use Gemini Vision; motion primitives (navigate/pick/place/push) via MoveIt. Four household tasks in three rooms, 25 trials each, 100 total:
| Task | Content | VAP-TAMP | OK-Robot | COWP | Closed-World |
|---|---|---|---|---|---|
| T1 Collect Cup | Fetch a cup across rooms | 92% | 84% | 80% | 40% |
| T2 Cut Lemon | Find lemon and knife, then cut (highest situation rate) | 84% | 72% | 64% | 24% |
| T3 Store Firewood | Collect wooden pieces into storage | 92% | 76% | 76% | 36% |
| T4 Pick Plate | Retrieve a plate from a cluttered table | 84% | 72% | 64% | 28% |
| Total | 100 trials | 88% | 76% | 69% | 33% |
The situation distribution is telling: 72 of 100 trials included at least one unforeseen situation (grasp failure 24, drop 18, scene change 16, visual 14). T2 has the highest rate due to its multi-step nature and frequent scene disturbances. Baselines span three paradigms: OK-Robot (open-vocabulary mobile manipulation SOTA, no active perception), COWP (LLM commonsense action-knowledge augmentation, task planning only, no motion planning), Closed-World (classical TAMP assuming perfect state estimation, no situation handling).
Viewpoint efficiency: 1.61 vs 4.35
Against greedy perception (exhaustively visiting predefined viewpoints around the target), greedy averages 4.35 viewpoints per disambiguation; VLM-guided selection averages 1.61 — a 63% reduction.
Failure modes: the bottleneck moved from perception to hardware
VAP-TAMP's 12 failures are 42% manipulation (hardware) errors with only 17% visual ambiguity, while OK-Robot (24 failures) and COWP (31) show 42%/39% visual ambiguity, and Closed-World's failures are 65% visual. The authors' summary is apt: active perception plus situation handling shifts the reliability bottleneck from lacking perceptual information to lacking physical capability.
Simulation ablations: why predicate verification beats "did it succeed"
OmniGibson simulation, five tasks with injected failure probabilities (grasp and cut at 50% combined failure rate each; open/close/turnon at 10%). Compared against five VLM verification strategies:
- VLM-planner / Classical-planner: no verification at all (avg 15%/25%);
- SuccessVQA: post-execution "did the action succeed?" (avg 52%);
- AffordanceVQA: pre-execution "is it possible?" (avg 35%);
- Suc.Aff.-QA: combining both (avg 67%) — still below predicate-based verification (avg 75%), because VLMs reason more reliably about observable state predicates than abstract outcomes.
Stage ablation: both stages 66.5% → effects only 53.0% → preconditions only 41.5%. Effect verification contributes more, since high-rate failures like grasp and cut are only detectable post-execution.
Execution-time curves and robustness
Beyond success rate, the execution-time plot (300 s timeout) shows VAP-TAMP reaching the highest success at the lowest execution times — active perception looks like "spending more time looking" but avoids repeated retries, so it is faster overall. Per task, baselines vary far more (OK-Robot 64–84%, Closed-World 24–40%) while VAP-TAMP holds 84–92%: situation handling not only raises the mean but compresses the variance — not collapsing on situation-heavy tasks is what long-term autonomy is built on.
What the head-to-head with OK-Robot shows
OK-Robot is the most instructive baseline: it also uses CLIP retrieval plus VLM situation handling, differing only in lacking active perception — one VLM judgment, final. The 12-point gap (88% vs 76%) is almost entirely attributable to that single change, giving a clean ablation: in the same open world with the same situations, "the ability to take another look" is worth 12 points. The COWP-versus-Closed-World gap (69% vs 33%) measures situation handling itself: classical TAMP is nearly unusable in an environment where 72/100 trials contain situations. Together the numbers quantify the thesis: active perception (+12) and situation handling (+36) are orthogonal, additive steps.
The failure-injection protocol
The simulation study also contributes a controllable failure-injection protocol: explicit per-action failure probabilities (grasp: 25% object unchanged + 25% dropped nearby; cut likewise; placein/placeon/open/close/turnon at 10%; fill at 5%; held object drops during navigation 10%). This isolises the comparison from random environmental noise — every method faces the identical failure distribution, so differences come purely from detection and recovery mechanisms. The protocol is directly reusable for follow-up work on execution monitoring.
Why this paper deserves a careful read
- It answers the most awkward question for VLM-in-robotics: what if it's wrong? Most systems treat VLMs as oracles; VAP-TAMP builds prompts from action knowledge, exposes hesitation through consistency checks, and resolves ambiguity by physically moving — VLM uncertainty becomes a first-class citizen of the planning loop.
- PDDL is not decoration: preconditions and effects are both the planner's input and the generator of VLM questions — symbolic knowledge and neural perception form a bidirectional loop, which is where "active" comes from.
- The experiments are designed with controls in mind: a 72/100 situation rate shows genuine difficulty; failure-mode analysis proves bottleneck shifting rather than vaguely reporting success rates; injected failure probabilities make ablations interpretable.
- For teams building service robots or open-world execution monitoring, the verify-execute-verify structure transfers directly.
Limitations and outlook
Three directions: whole-body optimization (base navigation and upper-body manipulation are currently optimized separately); reinforcement learning so the robot improves from failures; interactive perception that physically changes world configurations to gather information (e.g., pushing an occluder aside before verifying a predicate).
The framework also assumes the robot can localize in a known map and that target objects are detectable, with symbolic property inference resting entirely on the VLM; VLM hallucination accounts for 8% of failures — the lowest among all methods, but still worth compressing. As VLM reasoning improves, the same verify-execute-verify skeleton can upgrade to stronger models without structural changes — the dividend of decoupling verification logic from perception capability.
References
- Paper: arXiv:2604.26988 — Robot Planning and Situation Handling with Active Perception (IROS 2026; SUNY Binghamton / CMU / Ford Research / Agility Robotics)
- Project page: vap-tamp.github.io/vap-tamp
- Related systems: OK-Robot (RSS 2024), COWP (Autonomous Robots 2023), LLM-GROP (IJRR 2025), TAMPURA, PDDLStream / FFRob, SayCan, PaLM-E, VoxPoser
- Foundations: Integrated Task and Motion Planning (Annual Review 2021), Active Perception (Bajcsy, Proc. IEEE 1988)
SOURCE LINKS