Training-free
免训练CommonCompleting a new task by combining off-the-shelf models or algorithms, without updating any model parameters at all.
Training-free means a method applied to a new task makes no gradient updates whatsoever: it calls already-trained foundation models — large language models, vision-language models, segmentation or pose-estimation models — directly, chaining them together with prompting, code generation, search, or an optimization solver. Its appeal is not needing to collect robot data at all; switching tasks only means changing the instruction, which suits data-scarce embodied settings — its limitation is that performance is capped by the off-the-shelf models' own abilities and by how the intermediate representation is designed, and it usually struggles with fine, contact-rich motion. It isn't quite the same as zero-shot: zero-shot emphasizes not having seen samples of the target task, though the underlying model may well have been specially trained; training-free emphasizes that the whole method does no training at all. Papers also use the term for plug-and-play inference speedups, such as training-free visual-token pruning.
ExampleStanford's VoxPoser has a large language model write code that calls a vision-language model, turning instructions like “hang the towel on the rack” or “close the top drawer” into a 3D value map, which a motion planner then turns into a trajectory; the paper states explicitly that the whole pipeline involves no additional training.
- Also called
- No-training Method
- Related
- Zero-shot · Foundation Model · VoxPoser · ReKep · Code as Policies · Visual Token Pruning
- Sources
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models (project page)
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation (arXiv 2409.01652)