2:26Contact-Rich Humanoid Policies Trained Entirely Inside a World Model@chris_j_paxton · 308 views · 2026-09-04Coupled Local and Global World ModelsFoGFirst-Order RL
0:19Facet-0: a contact-aware foundation model for precise assembly@robotsdigest · 207 views · 2026-09-02VLAManuFacet-1KFacet-0
FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich ManipulationVision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but they are computationally expensive and fundamentally limited when temporal events are visually ambiguous, e.g., pushing a button multiple times with small movements. We propose FM-VLA, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, contact-rich manipulation. We encode force histories into compact force memory tokens with a variational autoencoder (VAE) pretrained with force time series reconstruction. By projecting force latent representations and short state history as additional conditioning tokens to the action expert module, we enable VLAs to leverage accumulated contact event history to guide manipulation. We evaluate FM-VLA on three memory-dependent tasks, including finding a hidden block, pressing a button, and wiping a dish for a specific number of times. Our lightweight force memory achieves over 80% success rate with minimal inference overhead, significantly outperforming baseline approaches. Project page: https://qft-333.github.io/FM-VLA-Page/Ruicheng Li, Qixiu Li, Ruichun Ma·Jul 20, 2026VLAVAEFM-VLAJul 20, 2026