Entropy Regularised Deterministic Optimal Control: From Path Integral Solution to Sample-Based Trajectory Optimisation

10/06/2021
by   Tom Lefebvre, et al.
0

Sample-based trajectory optimisers are a promising tool for the control of robotics with non-differentiable dynamics and cost functions. Contemporary approaches derive from a restricted subclass of stochastic optimal control where the optimal policy can be expressed in terms of an expectation over stochastic paths. By estimating the expectation with Monte Carlo sampling and reinterpreting the process as exploration noise, a stochastic search algorithm is obtained tailored to (deterministic) trajectory optimisation. For the purpose of future algorithmic development, it is essential to properly understand the underlying theoretical foundations that allow for a principled derivation of such methods. In this paper we make a connection between entropy regularisation in optimisation and deterministic optimal control. We then show that the optimal policy is given by a belief function rather than a deterministic function. The policy belief is governed by a Bayesian-type update where the likelihood can be expressed in terms of a conditional expectation over paths induced by a prior policy. Our theoretical investigation firmly roots sample based trajectory optimisation in the larger family of control as inference. It allows us to justify a number of heuristics that are common in the literature and motivate a number of new improvements that benefit convergence.

READ FULL TEXT

page 1

page 7

research
05/06/2022

Optimal Control as Variational Inference

In this article we address the stochastic and risk sensitive optimal con...
research
10/07/2019

Stochastic Optimal Control as Approximate Input Inference

Optimal control of stochastic nonlinear dynamical systems is a major cha...
research
09/13/2019

HJB Optimal Feedback Control with Deep Differential Value Functions and Action Constraints

Learning optimal feedback control laws capable of executing optimal traj...
research
03/23/2021

Smoothing-Averse Control: Covertness and Privacy from Smoothers

In this paper we investigate the problem of controlling a partially obse...
research
10/16/2020

Direct Policy Optimization using Deterministic Sampling and Collocation

We present an approach for approximately solving discrete-time stochasti...
research
03/01/2019

GuSTO: Guaranteed Sequential Trajectory Optimization via Sequential Convex Programming

Sequential Convex Programming (SCP) has recently seen a surge of interes...
research
01/08/2019

Solar-Sail Trajectory Design of Multiple Near Earth Asteroids Exploration Based on Deep Neural Network

In the preliminary trajectory design of the multi-target rendezvous prob...

Please sign up or login with your details

Forgot password? Click here to reset