DeepMind D4RT turns a video into a 4D world

Loading video
Loading videoGoogle DeepMind, UCL, and Oxford introduce D4RT, which jointly understands depth, motion, 3D correspondences, and camera parameters from a single video for 4D world models and embodied perception.
D4RT4D世界4D world世界模型world model空间AIspatial AI3D重建3D reconstruction视觉感知vision perception视频理解video understandingGoogle DeepMind具身智能embodied AIAlacritic_SuperTwitter
Category: perception
Author: @Alacritic_Super
Date: 2026-08-31T00:00:00
Duration: 213.379s
Reference: https://arxiv.org/abs/2512.08924





