Embodied AI Glossary中文

Inference Deployment

推理部署Essential

Getting a trained model running on real hardware or a server so it outputs actions in real time.

Inference deployment means moving a trained model from its training environment onto the hardware where it will actually be used; “inference” here means the model’s forward computation, not logical reasoning. In embodied AI, this means the model has to be wired up to real-time inputs — camera feeds, joint sensors — and output actions to the controller at a fixed rate. The hard parts are latency and compute: a robot’s onboard computer (such as a Jetson) has limited compute, so a large model often needs to be quantized, pruned, or accelerated with an engine such as TensorRT, or else run on a remote GPU server that sends actions back to the robot over the network (a policy server). How well the deployment is engineered directly affects whether the robot’s motion is smooth and whether it can correct itself in a closed loop.

ExampleA fine-tuned π0 model runs as a policy server on a GPU workstation; the robot periodically sends images and joint states over the network and receives back a chunk of actions.

Also called
Model Deployment, Deployment
Related
Inference Latency · On-Device / Edge Deployment · Policy Server (Remote Inference) · NVIDIA TensorRT · Post-Training Quantization · Inference
Sources
NVIDIA TensorRT
openpi (Physical Intelligence) - remote inference

See it in the full glossary →