Embodied AI Glossary中文

Molmo (Ai2)

MolmoAdvanced

An Ai2 vision-language model open-sourced with both weights and training data that can answer questions by “pointing” on the image.

Molmo is an open vision-language model family released by the Allen Institute for AI (Ai2) in September 2024, in several sizes — MolmoE-1B, Molmo-7B-O, Molmo-7B-D, and Molmo-72B — using CLIP as the vision encoder and OLMo or Qwen2, respectively, for the language component. It has two distinguishing features. First, its training dataset, PixMo, is released alongside it, and the data was not obtained by distilling a closed-source model — PixMo's detailed image captions come from annotators describing images out loud, then transcribed and cleaned up. Second, it can “point”: it can output the 2D coordinates of an object in an image directly, useful for counting and localization. Pointing is practical for robots, since it can tell a robot where to grasp or where to place something, and Ai2's later MolmoAct was built on top of Molmo. Molmo 2, released in December 2025, extended these abilities to video understanding, spatiotemporal grounding, and object tracking.

ExampleAsk Molmo “how many cups are in this image,” and it first places a point on each cup, then gives the total count.

Also called
Molmo and PixMo, Molmo 2
Related
Vision-Language Model · Pointing · MolmoAct · Open-weight Model · Allen Institute for AI · CLIP
Sources
Molmo (Ai2 blog)
Molmo and PixMo (arXiv 2409.17146)
Molmo 2 (Ai2 blog)
As of
2025-12

See it in the full glossary →