Dynamic Dialogue Policy Transformer for Continual Reinforcement Learning

04/12/2022
by   Christian Geishauser, et al.
0

Continual learning is one of the key components of human learning and a necessary requirement of artificial intelligence. As dialogue can potentially span infinitely many topics and tasks, a task-oriented dialogue system must have the capability to continually learn, dynamically adapting to new challenges while preserving the knowledge it already acquired. Despite the importance, continual reinforcement learning of the dialogue policy has remained largely unaddressed. The lack of a framework with training protocols, baseline models and suitable metrics, has so far hindered research in this direction. In this work we fill precisely this gap, enabling research in dialogue policy optimisation to go from static to dynamic learning. We provide a continual learning algorithm, baseline architectures and metrics for assessing continual learning models. Moreover, we propose the dynamic dialogue policy transformer (DDPT), a novel dynamic architecture that can integrate new knowledge seamlessly, is capable of handling large state spaces and obtains significant zero-shot performance when being exposed to unseen domains, without any growth in network parameter size.

READ FULL TEXT

page 13

page 14

page 15

research
12/31/2020

Continual Learning in Task-Oriented Dialogue Systems

Continual learning in task-oriented dialogue systems can allow us to add...
research
06/26/2020

Bookworm continual learning: beyond zero-shot learning and continual learning

We propose bookworm continual learning(BCL), a flexible setting where un...
research
05/23/2023

Continual Dialogue State Tracking via Example-Guided Question Answering

Dialogue systems are frequently updated to accommodate new services, but...
research
12/10/2019

How to Evaluate the Next System: Automatic Dialogue Evaluation from the Perspective of Continual Learning

Automatic dialogue evaluation plays a crucial role in open-domain dialog...
research
03/14/2022

L2Explorer: A Lifelong Reinforcement Learning Assessment Environment

Despite groundbreaking progress in reinforcement learning for robotics, ...
research
07/15/2022

How to Reuse and Compose Knowledge for a Lifetime of Tasks: A Survey on Continual Learning and Functional Composition

A major goal of artificial intelligence (AI) is to create an agent capab...
research
05/24/2023

Zero-shot Task Preference Addressing Enabled by Imprecise Bayesian Continual Learning

Like generic multi-task learning, continual learning has the nature of m...

Please sign up or login with your details

Forgot password? Click here to reset