Skild AI S1 Learns Unseen Tasks From a Single Video Demonstration

Loading video
Loading videoSkild AI S1 takes a recorded video of a task as its prompt, reads the demonstrated intent, objects and sequence through in-context learning, and maps them onto the robot in front of it with no weight updates and no task-specific post-training. It handles unfamiliar long-horizon tasks up to about 10 minutes, from plant potting and pancake making to pour-over coffee and kit assembly, with NVIDIA infrastructure spanning synthetic data, Isaac Lab, Cosmos and real-world deployment.
Robot foundation modelsSkild AIIn-Context LearningVideo DemonstrationsLong-Horizon ManipulationNVIDIA
Category: research
Author: @NVIDIARobotics
Date: 2026-09-10T00:00:00
Duration: 16.8s





