RT-SAFE: real-time embodied-agent safety benchmarking

Loading video
Loading videoRT-SAFE evaluates eight vision-language models while the world keeps moving: task success reaches 94.1%, but safe completion is only 0.7%, with 12.3x more collisions than static evaluation.
Embodied AISafety EvaluationReal-Time ReasoningVision-language modelsAutonomous NavigationCollision AvoidanceBenchmarksRT-SAFEReal-Time Inference
Category: research
Author: @XTRose88
Date: 2026-10-06T00:00:00
Duration: 60.0s
Reference: https://xtrose29.github.io/RT-SAFE/





