Embodied AI Glossary中文

CALVIN Benchmark

CALVINCommon

A tabletop manipulation benchmark testing whether a robot can follow 5 language instructions in a row.

CALVIN (Composing Actions from Language and Vision) is an open-source simulated benchmark released by Oier Mees, Wolfram Burgard, and colleagues at the University of Freiburg in Germany, published in RA-L 2022, where it won that journal's best-paper award that year. The setup is a table with a 7-DOF Franka arm, plus a drawer, a sliding door, a button, a switch, and three colored blocks, simulated in PyBullet; there are four environments, A, B, C, and D, structurally identical but differing in texture and part placement. It provides about 24 hours of teleoperated “play” data (only 1% of which is labeled with language), defines 34 task types, and its main evaluation requires executing 5 language instructions in a row, scored with Average Length. The most commonly used ABC→D setting trains on three environments and tests on the fourth, unseen one, to measure generalization.

ExampleOne test chain from the paper: “open the drawer” → “push the block into the drawer” → “take the block back out of the drawer” → “stack the blocks” → “close the drawer,” with the robot only advancing to the next instruction once it completes the current one.

Also called
CALVIN ABC→D, Composing Actions from Language and Vision
Related
Average Length (CALVIN) · Language-conditioned Policy · Long-horizon Task · LIBERO Benchmark · PyBullet · Play Data
Sources
CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (arXiv 2112.03227)
CALVIN GitHub 仓库 (Chinese)

See it in the full glossary →