Embodied AI Glossary中文

CLIP on Wheels

CoW(CLIP on Wheels)CoWAdvanced

A baseline that bolts an open-vocabulary model like CLIP onto a mobile robot to find objects from text, with no navigation training.

CoW was proposed by Shuran Song's group at Columbia University with the University of Washington in 2022, published at CVPR 2023. It studies “language-driven zero-shot object navigation”: a robot must find a target described in a sentence (such as “the toy airplane under the bed”) in an unfamiliar house, with no navigation training on those specific objects or scenes. CoW's approach is deliberately simple: while it's not yet confident it recognizes the target, it wanders using a classical exploration strategy; once an open-vocabulary model like CLIP (which can recognize objects from arbitrary text) locates the target in view with enough confidence, it plans a path there. The authors evaluated 21 CoW variants and proposed the Pasture benchmark, testing rare objects, objects described by appearance or spatial relation, and occluded objects. The best CoW beat the previous best method by 15.6 percentage points on a RoboTHOR object subset.

ExampleFor a goal like “tie-dye surfboard” — a category absent from navigation datasets — CoW first explores the room, and once CLIP judges some view a good enough match for that phrase, it plans a path to that location as the target.

Also called
CoW, CoWs on Pasture
Related
Object-Goal Navigation · CLIP · Open-vocabulary · Zero-shot · Frontier-based Exploration · VLFM
Sources
CoWs on Pasture (arXiv:2203.10421)
CoWs on Pasture 项目主页(CVPR 2023) (Chinese)
As of
2023-06

See it in the full glossary →