Embodied AI Glossary中文

RoboCerebra (A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation)

RoboCerebraAdvanced

A large simulation benchmark testing planning, reflection, and memory in long-horizon robot manipulation.

RoboCerebra was released in June 2025 by researchers at Beihang University, the National University of Singapore, Shanghai Jiao Tong University, and others, and was accepted to NeurIPS 2025. It focuses on “System 2”-style slow, deliberate reasoning: GPT is first used to generate long household tasks and break them into subtask sequences, which a human then carries out step by step inside simulation, yielding 100 task variants and 1,000 human-executed trajectories averaging about 2,972 simulation steps each — roughly 6 times longer, the paper says, than existing long-horizon manipulation datasets — with subtask time spans annotated. Test tasks fall into six categories: ideal conditions, random disturbance, observation inconsistency, memory exploration, memory execution, and mixed, with scenes changing mid-execution. A companion hierarchical framework uses a vision-language model (VLM) for high-level planning that writes to memory, and OpenVLA for low-level action execution; the paper uses this setup to compare how GPT-4o, Qwen2.5-VL, and other VLMs perform as planners on planning, reflection (judging whether a subtask is done), and memory.

ExampleGiven the instruction “prepare a drink, then tidy the table,” the high-level model must first break it into steps like fetching a cup, pouring the drink, and putting things away; if an object is randomly moved mid-execution, it also has to judge whether the current subtask is still complete and replan.

Related
Long-horizon Task · Dual-System Architecture (System 1 / System 2) · Embodied Memory · LIBERO Benchmark · OpenVLA · Benchmark
Sources
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation (arXiv 2506.06677)
As of
2025-10

See it in the full glossary →