Embodied AI Glossary中文

Policy Server (Remote Inference)

策略服务器Common

Running a large policy on a GPU server and having the robot send it observations and get actions back.

A policy server is a deployment pattern: a large model policy, such as a VLA, runs on a GPU-equipped workstation or in the cloud, while a lightweight client on the robot sends camera images, joint states, and instructions over the network; the server runs inference and returns a chunk of actions, which the client passes to the low-level controller for execution. It solves the problem of a robot's onboard compute not being enough to hold a model with billions of parameters, and it also makes it easy for multiple robots to share one model, or to swap models without touching the robot-side code. The tradeoff is network latency and jitter, so it's usually paired with action chunking and asynchronous inference, letting the robot execute its current action chunk while requesting the next one. openpi serves its policy over WebSocket, and LeRobot also has a gRPC-based asynchronous inference option.

ExampleStart openpi's serve_policy script on an RTX 4090 workstation to load π0; the client on the robot arm sends an observation each cycle and gets back a chunk of actions.

Also called
remote inference, server-client deployment, inference server
Related
Inference Deployment · Asynchronous Inference · Inference Latency · Action Chunking · openpi (Physical Intelligence) · Cloud-Edge-Device Collaboration
Sources
Physical-Intelligence/openpi - GitHub
LeRobot docs: Asynchronous Inference

See it in the full glossary →