HomeBody: a frontier VLM drives a Unitree G1 skill library directly, with no learned VLA in between

The poster highlights HomeBody from Stanford: a Unitree G1 first explores an unfamiliar kitchen, builds persistent spatial memory and a Real2Sim digital twin, and then uses GPT Astra (written as GPT-6 Astra in the tweet) to plan and compose skills. The project page comes from Stanford The Movement Lab with Caltech (September 2026; Gio Huh, Cayden Gu, Takara E. Truong, C. Karen Liu and Guy Tevet), and its central question is whether a learned System 1 VLA is still needed once the System 2 frontier VLM is strong enough. HomeBody replaces the VLM to VLA to System 0 whole-body controller chain with a plug-and-play VLM calling a composable skill library, using no environment-specific training data and no additional policy learning. Deployment has three steps: the robot explores an unseen kitchen while collecting 0.5x iPhone video, D435i observations, LiDAR scans with SLAM, joint poses and waypoints chosen by Astra; Astra then acts as the Real2Sim agent and rebuilds a digital twin in Isaac Sim from the data the robot collected itself; finally an everyday task is issued and the VLM chooses actions and targets from spatial context rather than following an action-level script. The skill library covers pick, place, open drawer, pick from drawer and navigate, all sharing one interface for targets and execution results, invoked through structured tool calls, and users can add their own learned policy or classical controller. The two demonstrated tasks are tidying a kitchen (gather the coffee bags on the island and discard the spoiled milk and orange juice cartons) and retrieving medicine that starts out of view, where stored keyframes locate the drawer, the right hand opens it, the medicine is handed to the person and the left hand discards a carton. For execution, Super Odometry localizes the G1 and ICP aligns the SLAM map with the reconstruction; a pick call supplies an image point normalized to 0-1000 plus which hand to use, which prompts segmentation with SAM 2.1 tracking and Fast-FoundationStereo depth, followed by minimum-jerk spline references, inverse kinematics and swept collision checks for the arm. Failures first trigger bounded local retries, and when recovery is exhausted the skill returns the reason to the VLM, which can reposition, pick a new target or change the plan. Lower-body control uses pretrained AMO, arm and hand commands run at 250 Hz with the AMO policy updated every fifth tick at 50 Hz, and the local stack runs on a single Razer Blade laptop with an RTX 4090 while Astra runs remotely. The authors also list limits: Real2Sim adds setup time and API cost, task length is constrained by reach, manipulation capability and hardware endurance including finger-servo overheating, and Astra reasoning latency introduces pauses between skills. The code is public on GitHub and the page does not yet link a paper.





