Show-Harness: Letting a VLM Agent Play Real Robots

Loading video
Loading videoShow-Harness is an embodied harness that exposes discrete semantic action units so VLMs can control real robots directly: no calibration and no extra VLA. Frontier closed-source models work zero-shot, and small open-source ones need only a few GPU-hours of fine-tuning.
VLAVLM AgentRobot ManipulationZero-ShotSemantic Action InterfaceEmbodied AIShow-HarnessShanghai AI LabVision-language models
Category: research
Author: @ZechenBai
Date: 2026-09-10T00:00:00
Duration: 110.033s
Reference: https://arxiv.org/abs/2609.10522





