Experience Replay Using Transition Sequences

Experience replay is one of the most commonly used approaches to improve the sample efficiency of reinforcement learning algorithms. In this work, we propose an approach to select and replay sequences of transitions in order to accelerate the learning of a reinforcement learning agent in an off-policy setting. In addition to selecting appropriate sequences, we also artificially construct transition sequences using information gathered from previous agent-environment interactions. These sequences, when replayed, allow value function information to trickle down to larger sections of the state/state-action space, thereby making the most of the agent's experience. We demonstrate our approach on modified versions of standard reinforcement learning tasks such as the mountain car and puddle world problems and empirically show that it enables better learning of value functions as compared to other forms of experience replay. Further, we briefly discuss some of the possible extensions to this work, as well as applications and situations where this approach could be particularly useful.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/18/2022

Neighborhood Mixup Experience Replay: Local Convex Interpolation for Improved Sample Efficiency in Continuous Control Tasks

Experience replay plays a crucial role in improving the sample efficienc...
research
09/13/2023

Attention Loss Adjusted Prioritized Experience Replay

Prioritized Experience Replay (PER) is a technical means of deep reinfor...
research
10/28/2022

Using Contrastive Samples for Identifying and Leveraging Possible Causal Relationships in Reinforcement Learning

A significant challenge in reinforcement learning is quantifying the com...
research
11/24/2020

Time Limits in Reinforcement Learning

In reinforcement learning, it is common to let an agent interact for a f...
research
09/26/2022

Paused Agent Replay Refresh

Reinforcement learning algorithms have become more complex since the inv...
research
06/12/2018

Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains

Model-based strategies for control are critical to obtain sample efficie...
research
02/15/2018

Prioritized Sweeping Neural DynaQ with Multiple Predecessors, and Hippocampal Replays

During sleep and awake rest, the hippocampus replays sequences of place ...

Please sign up or login with your details

Forgot password? Click here to reset