Code as Policies
代码即策略CaPCommonHas a large language model write Python code directly, calling perception and control APIs to command a robot.
Code as Policies was released by Robotics at Google in September 2022. Earlier work using LLMs to control robots, such as SayCan, mostly had the model pick the next step from a fixed list of skills. Code as Policies instead has an LLM — good at writing code — generate a Python program directly from a natural-language instruction: it calls perception APIs like object detection to get positions, uses NumPy to compute coordinates, then calls control primitives such as grasp or move, and can even write loops and conditionals. When it hits an undefined function, it recursively generates that function too, a process called hierarchical code generation. This lets it handle vague instructions that need spatial reasoning or concrete numeric values. The paper demonstrated this on tabletop manipulation, whiteboard drawing, and mobile robots, and the idea of an LLM writing code to drive a robot resurfaces later in VoxPoser and Eureka.
ExampleGiven the instruction “arrange the blocks in a horizontal line near the top,” the generated code first detects every block's position, computes a row of evenly spaced target points above the table, then calls the pick-and-place function for each one in turn.
- Also called
- CaP, Code as Policies: Language Model Programs for Embodied Control
- Related
- Large Language Model · SayCan · LLM-based Task Planning · VoxPoser · Eureka · Skill Primitive
- Sources
- Code as Policies (arXiv 2209.07753)
Code as Policies 项目主页 (Chinese) - As of
- 2022-09