Embodied AI Glossary中文

Vector Quantization

向量量化VQAdvanced

Preparing a 'codebook' and replacing a continuous vector with the ID of its nearest codeword, turning it into a discrete token.

Vector quantization is a classic compression technique from signal processing, systematically developed in the early 1980s by Robert Gray and colleagues. It works by preparing a set of representative vectors called a codebook, where each vector is called a codeword; for any input vector, you find the nearest codeword and record only its ID. This is essentially the same as k-means clustering, with codewords as the cluster centers. In deep learning, VQ is one of the main tools for turning continuous signals like images, audio, and actions into discrete tokens, which then makes it possible to model them with a language model's 'next-token prediction.' Residual vector quantization (RVQ) uses several codebooks in sequence, each quantizing the error left over from the previous level; Google's 2021 SoundStream audio codec uses it to let a single model switch bitrate anywhere between 3 and 18 kbps.

ExampleVQ-BeT (ICML 2024) encodes a robot's continuous actions into discrete tokens using hierarchical residual vector quantization, replacing the k-means clustering of its predecessor BeT, and the paper reports about 5x faster inference than Diffusion Policy.

Also called
VQ, Codebook, Residual VQ, RVQ
Related
Vector-Quantized Variational Autoencoder · Action Tokenizer · Finite Scalar Quantization · Tokenizer · Behavior Transformer · Latent Action
Sources
Wikipedia: Vector quantization
SoundStream: An End-to-End Neural Audio Codec (arXiv:2107.03312)
Behavior Generation with Latent Actions (VQ-BeT, arXiv:2403.03181)

See it in the full glossary →