Context-Based Soft Actor Critic for Environments with Non-stationary Dynamics

05/07/2021
by   Yuan Pu, et al.
0

The performance of deep reinforcement learning methods prone to degenerate when applied to environments with non-stationary dynamics. In this paper, we utilize the latent context recurrent encoders motivated by recent Meta-RL materials, and propose the Latent Context-based Soft Actor Critic (LC-SAC) method to address aforementioned issues. By minimizing the contrastive prediction loss function, the learned context variables capture the information of the environment dynamics and the recent behavior of the agent. Then combined with the soft policy iteration paradigm, the LC-SAC method alternates between soft policy evaluation and soft policy improvement until it converges to the optimal policy. Experimental results show that the performance of LC-SAC is significantly better than the SAC algorithm on the MetaWorld ML1 tasks whose dynamics changes drasticly among different episodes, and is comparable to SAC on the continuous control benchmark task MuJoCo whose dynamics changes slowly or doesn't change between different episodes. In addition, we also conduct relevant experiments to determine the impact of different hyperparameter settings on the performance of the LC-SAC algorithm and give the reasonable suggestions of hyperparameter setting.

READ FULL TEXT

page 5

page 7

page 8

research
04/17/2017

Pseudorehearsal in actor-critic agents

Catastrophic forgetting has a serious impact in reinforcement learning, ...
research
08/19/2022

Unified Policy Optimization for Continuous-action Reinforcement Learning in Non-stationary Tasks and Games

This paper addresses policy learning in non-stationary environments and ...
research
07/02/2019

Modified Actor-Critics

Robot Learning, from a control point of view, often involves continuous ...
research
04/13/2021

TASAC: Temporally Abstract Soft Actor-Critic for Continuous Control

We propose temporally abstract soft actor-critic (TASAC), an off-policy ...
research
10/05/2022

Real-Time Reinforcement Learning for Vision-Based Robotics Utilizing Local and Remote Computers

Real-time learning is crucial for robotic agents adapting to ever-changi...
research
07/29/2020

Learning Object-conditioned Exploration using Distributed Soft Actor Critic

Object navigation is defined as navigating to an object of a given label...
research
01/15/2022

Profitable Strategy Design by Using Deep Reinforcement Learning for Trades on Cryptocurrency Markets

Deep Reinforcement Learning solutions have been applied to different con...

Please sign up or login with your details

Forgot password? Click here to reset