MiniCPM-Robot series (ModelBest / OpenBMB)
面壁 MiniCPM-Robot 系列AdvancedMiniCPM's embodied model family, focused on small, on-device models for manipulation and for following a target on command.
This is an embodied model family open-sourced in July 2026 by the OpenBMB open-source community, led by ModelBest (面壁智能), with two models in the first release. MiniCPM-RobotManip is a 1.5-billion-parameter generalist VLA (vision-language-action model), with one set of weights covering multiple tasks; it uses streaming inference to keep past frames in context, retaining up to about 1 minute of visual memory, and reuses MiniCPM-V 4.6's visual-token compression, cutting each frame from 256 tokens down to 64. The company reports results on benchmarks like LIBERO, CALVIN, and RoboTwin 2.0 that come close to or beat the much larger π0.5. MiniCPM-RobotTrack is built on MiniCPM4-0.5B, about 900 million parameters, dedicated to following a target on a language instruction; it runs purely on vision on a Unitree Go2 robot dog's onboard compute, at more than 5 frames per second with about 180 milliseconds of end-to-end latency, which the company says is the best open-source result on EVT-Bench.
ExampleSay “follow that person” on a Unitree Go2 EDU, and RobotTrack follows along using only its onboard camera and local compute; official demos include riding an elevator and passing through an underground parking garage.
- Also called
- MiniCPM-RobotManip, MiniCPM-RobotTrack
- Related
- Vision-Language-Action Model · On-device Model · Embodied Visual Tracking · Visual Token Pruning · Unitree Go2 · π0.5
- Sources
- OpenBMB/MiniCPM-Robot (GitHub)
- As of
- 2026-07