SEAR: Benchmark for Self-Evolving LLM Agents on Robots

Loading video
Loading videoSEAR is a new benchmark and framework to rigorously evaluate how LLM agents self-evolve on robots, testing 7 frontier models on 330 sim tasks and real robots, finding Astra excels at live correction while Fable excels at code-first tool building.
Category: research
Author: @BangzhengL
Date: 2026-10-07T00:00:00
Duration: 20.5s
Reference: http://sear.bot





