Embodied AI Glossary中文

The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation

The ColosseumAdvanced

A benchmark applying 14 kinds of systematic environmental perturbations to RLBench tasks to test manipulation policies' generalization.

The Colosseum was proposed by Wilbert Pumacay, Jiafei Duan, Dieter Fox, and colleagues, and published at RSS 2024. Built on the PyRep simulation framework, it selects 20 of RLBench's 100 tasks, and each task can be perturbed along 14 dimensions: the color, texture, and size of both the manipulated object and static objects, light color, table color and texture, background texture, number of distractor objects, camera pose, and object friction and mass. The authors used it to test methods including PerAct, RVT, R3M, MVP, and VoxPoser, and found that a single perturbation alone drops success rate by 30%–50%, and applying several perturbations together drops it by more than 75%, with distractor count, target-object color, and lighting having the biggest effects. In a real-robot replication experiment, the R² between simulation and real-robot results was 0.614.

ExampleFor the same task, run separate trials changing only the table texture, only adding distractors, and only changing the camera pose, then compare each success rate against the unperturbed baseline to see which kind of change the policy is most sensitive to.

Also called
Colosseum
Related
RLBench · Generalization / Robustness Evaluation · Distractor Objects · Domain Randomization · PerAct · Visual Generalization
Sources
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation (arXiv 2402.08191)
The Colosseum 项目主页 (Chinese)
robot-colosseum GitHub 仓库 (Chinese)
As of
2024-05

See it in the full glossary →