Embodied AI Glossary中文

Language Grounding

语言接地Common

Connecting the words in language and instructions to actual objects, locations, and actions in the real world.

Language grounding means making sure the language a model understands actually corresponds to the physical world: which object on the table is “the red cup,” which region is “the left side,” which actions “open the drawer” requires. It traces back to the symbol grounding problem Stevan Harnad proposed in 1990: if a symbol is only ever defined in terms of other symbols, it never actually touches real meaning. Large language models have read enormous amounts of text but have never acted in a specific environment, so they can propose plans the robot in front of them actually cannot carry out. SayCan (Google and others, 2022) scores a language model's suggestions using each skill's value function, which estimates whether that skill can succeed in the current scene, grounding the plan in what the robot can actually do. Note that in Chinese, 落地 also commonly means commercial deployment, so context matters; the separate vision task of matching text to a location in an image is called visual grounding.

ExampleA user says, “I spilled my drink, help me.” SayCan has the language model list candidate steps, then uses the value function to pick the one that can actually succeed in the current scene, such as going to get a sponge first.

Also called
Grounding
Related
Symbol Grounding Problem · SayCan · Visual Grounding · Affordance · Instruction Following · Large Language Model
Sources
Symbol grounding problem - Wikipedia
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan, arXiv 2204.01691)

See it in the full glossary →