Bayesian Optimization with Automatic Prior Selection for Data-Efficient Direct Policy Search

09/20/2017
by   Rémi Pautrat, et al.
0

One of the most interesting features of Bayesian optimization for direct policy search is that it can leverage priors (e.g., from simulation or from previous tasks) to accelerate learning on a robot. In this paper, we are interested in situations for which several priors exist but we do not know in advance which one fits best the current situation. We tackle this problem by introducing a novel acquisition function, called Most Likely Expected Improvement (MLEI), that combines the likelihood of the priors and the expected improvement. We evaluate this new acquisition function on a transfer learning task for a 5-DOF planar arm and on a possibly damaged, 6-legged robot that has to learn to walk on flat ground and on stairs, with priors corresponding to different stairs and different kinds of damages. Our results show that MLEI effectively identifies and exploits the priors, even when there is no obvious match between the current situations and the priors.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
06/25/2020

Prior-guided Bayesian Optimization

While Bayesian Optimization (BO) is a very popular method for optimizing...
research
03/02/2020

Robust Policy Search for Robot Navigation with Stochastic Meta-Policies

Bayesian optimization is an efficient nonlinear optimization method wher...
research
08/02/2022

Learning Skill-based Industrial Robot Tasks with User Priors

Robot skills systems are meant to reduce robot setup time for new manufa...
research
12/20/2022

HyperBO+: Pre-training a universal prior for Bayesian optimization with hierarchical Gaussian processes

Bayesian optimization (BO), while proved highly effective for many black...
research
09/20/2017

Using Parameterized Black-Box Priors to Scale Up Model-Based Policy Search for Robotics

The most data-efficient algorithms for reinforcement learning in robotic...
research
07/06/2018

A survey on policy search algorithms for learning robot controllers in a handful of trials

Most policy search algorithms require thousands of training episodes to ...
research
07/16/2019

Adaptive Prior Selection for Repertoire-based Online Learning in Robotics

Among the data-efficient approaches for online adaptation in robotics (m...

Please sign up or login with your details

Forgot password? Click here to reset