Skill Primitive
原子技能CommonThe smallest reusable unit of action, such as grasp, place, or open drawer, combined to complete complex tasks.
This approach breaks a robot's abilities into a set of basic actions that can be called and trained separately — “grasp,” “place,” “push,” “open drawer,” “move to a location.” Each skill primitive usually takes parameters, such as which object to grasp or where to place it, and is called in sequence by a higher-level planner or large model to assemble a long-horizon task. The benefit is that each skill is easy to train and validate on its own, and can be reused across tasks; the cost is that anything outside the skill library simply cannot be done, and the transitions between skills are a common source of errors. A landmark example is Google's SayCan (2022), which has a large language model pick the next step from a pretrained skill library and uses each skill's value function to judge whether it is feasible right now; RAPS (2021) instead treats hand-defined, parameterized primitives as the action space for reinforcement learning, to improve exploration and learning efficiency.
Example“Put the soda in the fridge” can be broken into: navigate to the table, grasp the soda, navigate to the fridge, open the fridge door, place the soda, close the door — each step is a skill primitive.
- Also called
- Action Primitive
- Related
- Long-horizon Task · Hierarchical Architecture · Motion Primitives · SayCan · LLM-based Task Planning · Dynamic Movement Primitives
- Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan)
Accelerating Robotic Reinforcement Learning via Parameterized Action Primitives (RAPS)