Embodied AI Glossary中文

MIKASA-Robo

Advanced

A tabletop-manipulation benchmark specifically testing robot memory, with tasks requiring recall of information that's occluded or gone.

MIKASA-Robo is a memory-intensive manipulation benchmark proposed by Cherepanov, Panov, and colleagues in February 2025, part of MIKASA (a suite for evaluating memory-intensive skills), with its paper published at ICLR 2026. Many real-world tasks are partially observable: an object gets blocked from view, critical information appears only briefly, and a policy that only looks at the current frame has no way to succeed. It's built on ManiSkill3, with the first version containing 32 tasks. It later expanded into MIKASA-Robo-VLA, aimed at VLA models: the task count grew to 90, covering 10 categories of memory, each task paired with a language instruction, with 22,500 trajectories released on Hugging Face, usable to test memory-equipped models such as MemoryVLA.

ExampleIn the ShellGameTouch task, the robot can see the red ball under one of three positions for the first 5 steps; all three are then covered with cups, and the robot must touch the cup hiding the ball.

Also called
MIKASA-Robo-VLA
Related
Embodied Memory · Memory-Augmented VLA · Partially Observable Markov Decision Process · ManiSkill · Memory-Augmented VLA · Benchmark
Sources
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning (arXiv 2502.10550)
MIKASA-Robo GitHub 仓库 (Chinese)
MIKASA-Robo-VLA Documentation
As of
2026-09

See it in the full glossary →