Embodied AI Glossary中文

Finite Scalar Quantization

有限标量量化FSQAdvanced

Rounding each dimension of a continuous vector directly to a few fixed levels, replacing a VQ codebook for discretization.

Finite scalar quantization was proposed in 2023 by Mentzer and colleagues at Google Research, in a paper subtitled 'VQ-VAE Made Simple.' Traditional vector quantization (VQ) maintains a learnable codebook, and training often suffers from 'codebook collapse,' where large numbers of code entries never get used, requiring extra tricks like commitment loss and entropy penalties to fix. FSQ instead projects features down to very few dimensions (usually fewer than 10), clips each dimension to a range, and rounds it directly to a small number of fixed levels; the combination of levels across dimensions forms an implicit codebook. There's no codebook to learn and no collapse to worry about, and it performs on par with VQ on tasks like image generation and depth estimation. It's commonly used to turn images, video, or action sequences into discrete tokens for an autoregressive Transformer to process.

ExampleNVIDIA's Cosmos Tokenizer discrete version uses FSQ to compress video into discrete tokens, with an index range of about 64,000 (64K).

Also called
FSQ
Related
Vector Quantization · Vector-Quantized Variational Autoencoder · Action Tokenizer · Video Tokenizer · Token · NVIDIA Cosmos
Sources
Finite Scalar Quantization: VQ-VAE Made Simple (Mentzer et al., arXiv 2309.15505)
NVIDIA Cosmos Tokenizer (GitHub)

See it in the full glossary →