End-to-End
端到端E2EEssentialUsing one model to go straight from raw sensor input to control commands, with no hand-designed intermediate modules.
A traditional robot system is a modular pipeline: a perception module identifies objects and poses, a planning module computes a path, and a control module tracks the trajectory, each designed and tuned separately and passing information through human-defined interfaces such as object coordinates. End-to-end instead hands the whole chain to one neural network, mapping raw input, such as camera images and proprioceptive state, directly to joint positions or torque commands, trained all at once toward a single objective. An early landmark in robotics is Levine, Finn, and colleagues' deep visuomotor policy work (2015 preprint, published 2016), which used a convolutional network of only about 92,000 parameters to map raw camera images directly to motor torques. The benefit is less hand-engineering and no error compounding across modules, with performance able to grow with more data; the cost is needing a lot of data, and errors are harder to trace to a cause. Most VLAs today are described as end-to-end models.
ExampleLevine and colleagues had a PR2 robot's policy network read raw camera images directly, plus the robot's own joint angles and other state, and output torque commands for every joint, learning tasks such as screwing on a bottle cap and hanging a coat hanger on a rod with no separate object-detection or pose-estimation module at test time.
- Also called
- E2E, End-to-End Model, End-to-End Learning
- Related
- Visuomotor Policy · Vision-Language-Action Model · Hierarchical Architecture · Sense-Plan-Act · Imitation Learning · End-to-End Training of Deep Visuomotor Policies
- Sources
- End-to-End Training of Deep Visuomotor Policies (arXiv 1504.00702)