Embodied AI Glossary中文

SayCan

Common

Combines what a language model says is useful with what a value function says is achievable to pick each action.

SayCan is work released by the Google Robotics team and Everyday Robots in April 2022. Large language models have broad commonsense knowledge, but don't know what the robot in front of them can actually do in its current situation; letting an LLM write a plan directly often produces steps that aren't achievable. SayCan instead equips the robot with a set of pretrained skills (such as “pick up the sponge” or “go to the table”); at each step, the language model scores how useful each skill would be toward completing the instruction, and a value function learned through reinforcement learning scores how likely each skill is to succeed from the current state (its affordance); the two scores are multiplied together, and the highest-scoring skill is executed, repeating until the task ends. Using PaLM in place of the original language model, PaLM-SayCan reached an 84% planning success rate and a 74% execution success rate across 101 instructions in a real kitchen. It's one of the founding works on using large models for high-level robot task planning.

ExampleWhen a user says “I spilled my Coke, can you help me get something to clean it up,” SayCan selects, in sequence: find a sponge, pick up the sponge, bring it to you, done.

Also called
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, PaLM-SayCan
Related
Affordance · Language Grounding · LLM-based Task Planning · Inner Monologue · PaLM-E · Value Function
Sources
SayCan 项目主页 (Chinese)
As of
2022-08

See it in the full glossary →