GemBench
AdvancedAn RLBench-based simulation benchmark that tests a language-conditioned manipulation policy's generalization across four difficulty levels.
GemBench was proposed by Ricardo Garcia, Shizhe Chen, and Cordelia Schmid at Inria and ENS Paris, published at ICRA 2025. Built on the RLBench simulator, it defines 7 action primitives — press, grasp, push, turn, close, open, and place/stack — trains on 16 tasks (31 variants), and tests on 44 tasks (92 variants), with generalization difficulty split into four levels: new object placements, new rigid objects, new articulated objects, and new long-horizon tasks. The same paper also proposes 3D-LOTUS (a point-cloud-based language-conditioned policy) and 3D-LOTUS++ (which adds a large language model for task planning and a vision-language model for object localization). Results show that pure imitation-learning policies score near-perfectly on familiar tasks but drop off sharply when faced with unlearned combinations of skills.
ExampleAt level 1 (only object placement changes), 3D-LOTUS scores 94.3% success; at level 4, which requires combining learned actions into new long-horizon tasks, it scores only 0.3%, while 3D-LOTUS++ with added LLM planning reaches 17.4%.
- Also called
- GEMBench
- Related
- RLBench · Generalization · Compositional Generalization · Long-horizon Task · Language-conditioned Policy · Benchmark
- Sources
- Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy (arXiv 2410.01345)
GemBench project page - As of
- 2025-05