Embodied AI Glossary中文

Differentiable Simulation

可微仿真Common

A simulator whose output can be differentiated, so gradient descent can optimize actions or physical parameters directly.

An ordinary simulator only computes forward: given this action, what happens next. A differentiable simulator can also use automatic differentiation to compute, backward, the gradient of the result with respect to the action, the initial state, or physical parameters such as mass and friction. That makes it possible to directly optimize a control sequence, a policy network, or an estimate of physical parameters using gradient descent, the same way a neural network is trained, instead of estimating a gradient indirectly through large amounts of trial and error the way reinforcement learning does. Notable examples include DiffTaichi (ICLR 2020), Google's Brax, and MJX, the JAX version of MuJoCo (its Warp backend does not support automatic differentiation). The hard part is contact: collisions introduce sudden changes in the dynamics, so gradients can be huge, noisy, or zero; a 2022 ICML study by Suh and colleagues found that stiffness and discontinuity undermine the usefulness of this kind of first-order gradient.

ExampleThe DiffTaichi paper implements 10 differentiable simulators, and using them to optimize neural-network controllers typically converges within a few dozen iterations.

Also called
differentiable physics, differentiable physics engine
Related
Physics Engine · MuJoCo XLA · Brax · NVIDIA Warp · Trajectory Optimization · System Identification
Sources
DiffTaichi: Differentiable Programming for Physical Simulation (arXiv 1910.00935)
A Review of Differentiable Simulators (arXiv 2407.05560)
Do Differentiable Simulators Give Better Policy Gradients? (arXiv 2202.00817)

See it in the full glossary →