Embodied Safety
具身安全CommonEnsuring a robot acting in the real world does not injure people, damage property, or get hijacked by malicious commands.
Embodied safety concerns the risks that come from an embodied agent acting in the physical world, and breaks down roughly into two layers. One is traditional physical safety: collision detection, force and speed limits, emergency stops, corresponding to standards for collaborative robots such as ISO/TS 15066. The other is a newer set of problems that arise once large models are put into robots, which Google DeepMind calls semantic safety: the model may produce a dangerous action because of hallucination, prompt injection, or a commonsense mistake. The BadRobot paper (ICLR 2025) demonstrated that a voice-based jailbreak could make a large-model-based robot carry out harmful actions; DeepMind's 2025 work proposed the ASIMOV benchmark and an automatically generated “robot constitution” to evaluate and constrain model behavior.
ExampleA user tells a home robot, “pour this cleaning fluid into a cup and hand it to the guest.” Embodied safety requires the upper-level model to recognize this as a dangerous request and refuse it; the lower-level controller, in turn, must ensure that even if the arm makes a mistake, it never hits a person with excessive force.
- Also called
- Embodied AI Safety
- Related
- Functional Safety · Physical Human-Robot Interaction · Safe Reinforcement Learning · Control Barrier Function · Asimov's Three Laws of Robotics · ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots
- Sources
- Generating Robot Constitutions & Benchmarks for Semantic Safety (arXiv 2503.08663)
BadRobot: Jailbreaking Embodied LLM Agents in the Physical World (arXiv 2407.20242) - As of
- 2025-03