Auto-labeling
自动标注AdvancedUsing pretrained models or scripts, instead of humans, to automatically add language instructions, bounding boxes, and other labels to robot data.
Auto-labeling means using a program or a pretrained model — a vision-language model, an object detector, a large language model — to automatically add labels to collected data instead of a human doing it. Robot data often lacks several kinds of information: what a trajectory is doing (a language instruction), where it splits into subtasks, and where objects are in the frame (bounding boxes, masks). Manual annotation one item at a time is expensive and slow, and can't keep up once data volume grows. Two typical approaches: using an image-text model like CLIP to match un-annotated demonstrations to instructions, as Google's DIAL does; or, like Berkeley's ECoT, using Grounding DINO to box objects and a large model to write out task decomposition and reasoning steps as extra supervision for training a VLA with reasoning ability. Auto-labeling is less reliable than human annotation, so it is usually spot-checked or filtered with rules, and commonly paired with data quality inspection, hindsight relabeling, and instruction augmentation.
ExampleGoogle's 2022 DIAL used CLIP to automatically add language labels to about 80,000 demonstrations (96.5% of which originally had no crowd-sourced label), and the resulting policy could execute 60 new instructions absent from the original data.
- Also called
- Automated Annotation
- Related
- Data Annotation · Language Annotation · Hindsight Relabeling · Instruction Augmentation · Subtask Segmentation · Data Quality Control
- Sources
- Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models (DIAL, arXiv 2211.11736)
Robotic Control via Embodied Chain-of-Thought Reasoning (ECoT, arXiv 2407.08693)