Differential Dynamic Programming
微分动态规划DDPAdvancedA trajectory-optimization method that repeatedly makes a second-order approximation along the current trajectory and sweeps back and forth to improve the control sequence.
DDP is an iterative algorithm for solving nonlinear optimal control, proposed by David Mayne in 1966 and later systematized in a book of the same name with Jacobson. Given an initial control sequence, it repeats two steps: a backward pass, which takes a second-order expansion of the dynamics and cost around the current trajectory and works backward from the endpoint to compute a correction k = −Q_uu⁻¹Q_u and a feedback gain K = −Q_uu⁻¹Q_ux at each step, where Q is the total cost of ‘choosing control u this step and then acting optimally from then on,’ with subscripts denoting derivatives with respect to u or the state x; and a forward pass, which re-simulates a new trajectory using u = ū + αk + K(x − x̄) (ū, x̄ being the old trajectory and α a line-search step size). It converges quadratically near the optimum and produces feedback gains as a byproduct, making it well suited to MPC. Dropping the dynamics' second-derivative terms gives the more commonly used iLQR. DDP belongs to the shooting-method family, with the state obtained by simulation, which handles state constraints less conveniently than collocation methods.
ExampleThe open-source library Crocoddyl uses DDP and its variant FDDP as its core solvers, leveraging Pinocchio's analytical derivatives to compute optimal trajectories — and the accompanying feedback gains — with contact sequences for legged robots and similar systems.
- Also called
- DDP
- Related
- Iterative Linear Quadratic Regulator · Linear Quadratic Regulator · Trajectory Optimization · Model Predictive Control · Optimal Control · Crocoddyl (Contact RObot COntrol by Differential DYnamic programming Library)
- Sources
- Wikipedia: Differential dynamic programming
loco-3d/crocoddyl(solvers based on DDP algorithms)