Embodied AI Glossary中文

Embedding

嵌入向量Common

Representing a word, image patch, or action as a string of real numbers, where similar things end up as nearby vectors.

An embedding is the basic unit a neural network uses to represent information internally: a word, an image patch, a frame of robot state, or some other object gets mapped into a fixed-length string of real numbers, say 1,024 of them. It is learned during training, not hand-specified, and once trained, objects with similar meaning end up close together in vector space. Compared with one-hot encoding, where each category occupies its own dimension and only one entry is ever 1, embeddings use far fewer dimensions and can still express similarity; word2vec is an early, well-known example. Embeddings show up almost everywhere in embodied AI models: a tokenizer cuts text into tokens and looks up their word embeddings, a vision encoder turns each image patch into a visual-token embedding, and a projection layer then maps those into the language model's dimension; robot state and noisy actions also pass through a linear layer to become embeddings before joining other tokens into a Transformer.

ExampleCLIP encodes a photo of a cat and the sentence “a photo of a cat” into embedding vectors separately, and the cosine similarity between them comes out clearly higher than between that same photo and “a photo of a dog.”

Also called
Vector Representation
Related
Token · Tokenizer · Visual Token · Latent Space · Projector / Connector · Representation Learning
Sources
Embeddings | Machine Learning Crash Course (Google for Developers)
Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv 2103.00020)

See it in the full glossary →