Diluted Near-Optimal Expert Demonstrations for Guiding Dialogue Stochastic Policy Optimisation

11/25/2020
by   Thibault Cordier, et al.
0

A learning dialogue agent can infer its behaviour from interactions with the users. These interactions can be taken from either human-to-human or human-machine conversations. However, human interactions are scarce and costly, making learning from few interactions essential. One solution to speedup the learning process is to guide the agent's exploration with the help of an expert. We present in this paper several imitation learning strategies for dialogue policy where the guiding expert is a near-optimal handcrafted policy. We incorporate these strategies with state-of-the-art reinforcement learning methods based on Q-learning and actor-critic. We notably propose a randomised exploration policy which allows for a seamless hybridisation of the learned policy and the expert. Our experiments show that our hybridisation strategy outperforms several baselines, and that it can accelerate the learning when facing real humans.

READ FULL TEXT
POST COMMENT

Comments

There are no comments yet.

Authors

01/31/2018

Pretraining Deep Actor-Critic Reinforcement Learning Algorithms With Expert Demonstrations

Pretraining with expert demonstrations have been found useful in speedin...
07/24/2019

Efficient Exploration with Self-Imitation Learning via Trajectory-Conditioned Policy

This paper proposes a method for learning a trajectory-conditioned polic...
10/31/2017

Adversarial Advantage Actor-Critic Model for Task-Completion Dialogue Policy Learning

This paper presents a new method --- adversarial advantage actor-critic ...
09/16/2018

Inspiration Learning through Preferences

Current imitation learning techniques are too restrictive because they r...
06/14/2018

Self-Imitation Learning

This paper proposes Self-Imitation Learning (SIL), a simple off-policy a...
07/01/2020

Policy Improvement from Multiple Experts

Despite its promise, reinforcement learning's real-world adoption has be...
07/01/2020

Fighting Failures with FIRE: Failure Identification to Reduce Expert Burden in Intervention-Based Learning

Supervised imitation learning, also known as behavior cloning, suffers f...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.