Physical Intelligence (Sergey Levine's team, Mar 19, 2026) present RL Tokens (RLT): freeze a pretrained VLA, attach an encoder-decoder that compresses its internal embeddings into a bottleneck RL token, then run online RL with a ~1M-parameter actor-critic on the real robot, refining only the critical phase (editing VLA action chunks rather than generating from scratch, anchored by a BC regularizer plus reference-action dropout). Across four sub-millimeter tasks the critical phase speeds up by up to 3x, screw success goes 20% to 65%, and half of Ethernet insertion episodes beat every human teleoperation demo - with just 15 minutes of real robot data.
BLOG