Recurrent Network-based Deterministic Policy Gradient for Solving Bipedal Walking Challenge on Rugged Terrains

10/08/2017
by   Doo Re Song, et al.
0

This paper presents the learning algorithm based on the Recurrent Network-based Deterministic Policy Gradient. The Long-Short Term Memory is utilized to enable the Partially Observed Markov Decision Process framework. The novelty are improvements of LSTM networks: update of multi-step temporal difference, removal of backpropagation through time on actor, initialisation of hidden state using past trajectory scanning, and injection of external experiences learned by other agents. Our methods benefit the reinforcement learning agent on inferring the desirable action by referring the trajectories of both past observations and actions. The proposed algorithm was implemented to solve the Bipedal-Walker challenge in OpenAI virtual environment where only partial state information is available. The validation on the extremely rugged terrain demonstrates the effectiveness of the proposed algorithm by achieving a new record of highest rewards in the challenge. The autonomous behaviors generated by our agent are highly adaptive to a variety of obstacles as shown in the simulation results.

READ FULL TEXT
research
11/15/2019

Improved Exploration through Latent Trajectory Optimization in Deep Deterministic Policy Gradient

Model-free reinforcement learning algorithms such as Deep Deterministic ...
research
02/06/2021

Learning adaptive differential evolution algorithm from optimization experiences by policy gradient

Differential evolution is one of the most prestigious population-based s...
research
12/03/2021

Episodic Policy Gradient Training

We introduce a novel training procedure for policy gradient methods wher...
research
07/19/2023

Joint Service Caching, Communication and Computing Resource Allocation in Collaborative MEC Systems: A DRL-based Two-timescale Approach

Meeting the strict Quality of Service (QoS) requirements of terminals ha...
research
07/10/2018

Generalized deterministic policy gradient algorithms

We study a setting of reinforcement learning, where the state transition...
research
09/24/2018

EpiRL: A Reinforcement Learning Agent to Facilitate Epistasis Detection

Epistasis (gene-gene interaction) is crucial to predicting genetic disea...
research
07/11/2023

Safe Reinforcement Learning for Strategic Bidding of Virtual Power Plants in Day-Ahead Markets

This paper presents a novel safe reinforcement learning algorithm for st...

Please sign up or login with your details

Forgot password? Click here to reset