Nonlinear stochastic trajectory optimization for centroidal momentum motion generation of legged robots
Value learning from trajectory optimization and Sobolev descent A step toward reinforcement learning with superlinear convergence properties