Embodied AI Glossary中文

On-device Model

端侧模型Common

A model that runs locally on a device's own chip — a robot, a phone — instead of relying on the cloud.

An on-device model is deployed to run locally on an end device — for example, a robot's onboard computer (such as an NVIDIA Jetson), a phone chip, or a car's chip — as opposed to a model that runs inference in a cloud data center. Device compute, memory, and power are all limited, so on-device models are usually smaller and compressed using quantization, pruning, distillation, and similar techniques; Apple's 2025 on-device foundation model, for example, has about 3B parameters, with weights compressed to 2 bits each via quantization-aware training. For robots, running locally has real practical benefits: it still works without a network connection, avoids network round-trip latency, and keeps data from leaving the device. The tradeoff is that on-device models are usually less capable than large cloud models, so a common setup has a large cloud model handle slower planning and reasoning while a small on-device model outputs actions at high frequency — cloud-edge-device collaboration.

ExampleGoogle DeepMind released Gemini Robotics On-Device in June 2025: this VLA can run locally on the robot with no network connection, adapts to new tasks with just 50 to 100 demonstrations, and has been validated on a bimanual Franka FR3 and the Apptronik Apollo humanoid.

Also called
Edge Model, On-device AI
Related
On-Device / Edge Deployment · Cloud-Edge-Device Collaboration · Post-Training Quantization · NVIDIA Jetson · Gemini Robotics On-Device · Inference Latency
Sources
Google DeepMind: Gemini Robotics On-Device brings AI to local robotic devices
Apple Machine Learning Research: Updates to Apple's On-Device and Server Foundation Language Models
As of
2025-07

See it in the full glossary →