Bayesian Nonparametrics for Offline Skill Discovery

02/09/2022
by   Valentin Villecroze, et al.
0

Skills or low-level policies in reinforcement learning are temporally extended actions that can speed up learning and enable complex behaviours. Recent work in offline reinforcement learning and imitation learning has proposed several techniques for skill discovery from a set of expert trajectories. While these methods are promising, the number K of skills to discover is always a fixed hyperparameter, which requires either prior knowledge about the environment or an additional parameter search to tune it. We first propose a method for offline learning of options (a particular skill framework) exploiting advances in variational inference and continuous relaxations. We then highlight an unexplored connection between Bayesian nonparametrics and offline skill discovery, and show how to obtain a nonparametric version of our model. This version is tractable thanks to a carefully structured approximate posterior with a dynamically-changing number of options, removing the need to specify K. We also show how our nonparametric extension can be applied in other skill frameworks, and empirically demonstrate that our method can outperform state-of-the-art offline skill learning algorithms across a variety of environments. Our code is available at https://github.com/layer6ai-labs/BNPO .

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/20/2022

Learning and Retrieval from Prior Data for Skill-based Imitation Learning

Imitation learning offers a promising path for robots to learn general-p...
research
01/31/2023

Skill Decision Transformer

Recent work has shown that Large Language Models (LLMs) can be incredibl...
research
07/21/2023

Diverse Offline Imitation via Fenchel Duality

There has been significant recent progress in the area of unsupervised s...
research
06/14/2023

Skill-Critic: Refining Learned Skills for Reinforcement Learning

Hierarchical reinforcement learning (RL) can accelerate long-horizon dec...
research
10/15/2017

DDCO: Discovery of Deep Continuous Options for Robot Learning from Demonstrations

An option is a short-term skill consisting of a control policy for a spe...
research
01/19/2020

Learning Options from Demonstration using Skill Segmentation

We present a method for learning options from segmented demonstration tr...
research
02/18/2022

A practical DMPs Implementation for Skill Creation and Teleoperation with Assistive Manipulators

Assistive robotic manipulators are becoming increasingly important for p...

Please sign up or login with your details

Forgot password? Click here to reset