Embodied AI Glossary中文

Encoder-Decoder

编码器-解码器Common

A network design that compresses the input into an intermediate representation, then generates the output from that representation.

Encoder-decoder is a general-purpose neural network design: an encoder turns the input (a sentence, an image, a set of sensor readings) into an intermediate representation, and a decoder generates the output from that representation, with input and output allowed to differ in length. Sutskever et al.'s 2014 machine-translation model, which used two LSTMs (a type of recurrent neural network) to translate English to French, is the classic example on sequence-to-sequence tasks. The original 2017 Transformer also used this design, with the decoder reading the encoder's output through cross-attention. Later, GPT-style large language models mostly switched to decoder-only designs, but encoder-decoder remains common in robotics: ACT uses a Transformer encoder to fuse multiple camera views and joint state, then has a decoder output an entire chunk of actions at once. Autoencoders and variational autoencoders also belong to this family.

ExampleControlling the bimanual ALOHA robot, ACT's encoder takes in 4 camera views and a 14-dimensional joint-angle reading; its decoder outputs the target 14-dimensional joint positions for the next k steps (typically 100).

Also called
Encoder-Decoder Architecture, Seq2Seq, Sequence-to-Sequence Model
Related
Transformer · Decoder-only Architecture · Cross-Attention · Autoencoder · Variational Autoencoder · Action Chunking with Transformers
Sources
Dive into Deep Learning: The Encoder–Decoder Architecture
Sequence to Sequence Learning with Neural Networks (arXiv 1409.3215)
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)

See it in the full glossary →