Modeling Survival in model-based Reinforcement Learning

04/18/2020
by   SM, et al.
0

Although recent model-free reinforcement learning algorithms have been shown to be capable of mastering complicated decision-making tasks, the sample complexity of these methods has remained a hurdle to utilizing them in many real-world applications. In this regard, model-based reinforcement learning proposes some remedies. Yet, inherently, model-based methods are more computationally expensive and susceptible to sub-optimality. One reason is that model-generated data are always less accurate than real data, and this often leads to inaccurate transition and reward function models. With the aim to mitigate this problem, this work presents the notion of survival by discussing cases in which the agent's goal is to survive and its analogy to maximizing the expected rewards. To that end, a substitute model for the reward function approximator is introduced that learns to avoid terminal states rather than to maximize accumulated rewards from safe states. Focusing on terminal states, as a small fraction of state-space, reduces the training effort drastically. Next, a model-based reinforcement learning method is proposed (Survive) to train an agent to avoid dangerous states through a safety map model built upon temporal credit assignment in the vicinity of terminal states. Finally, the performance of the presented algorithm is investigated, along with a comparison between the proposed and current methods.

READ FULL TEXT
research
02/15/2022

Safe Reinforcement Learning by Imagining the Near Future

Safe reinforcement learning is a promising path toward applying reinforc...
research
03/20/2023

Deceptive Reinforcement Learning in Model-Free Domains

This paper investigates deceptive reinforcement learning for privacy pre...
research
06/20/2022

Guided Safe Shooting: model based reinforcement learning with safety constraints

In the last decade, reinforcement learning successfully solved complex c...
research
06/18/2016

On Reward Function for Survival

Obtaining a survival strategy (policy) is one of the fundamental problem...
research
09/06/2021

Method for making multi-attribute decisions in wargames by combining intuitionistic fuzzy numbers with reinforcement learning

Researchers are increasingly focusing on intelligent games as a hot rese...
research
03/13/2023

Transformer-based World Models Are Happy With 100k Interactions

Deep neural networks have been successful in many reinforcement learning...
research
03/01/2022

DreamingV2: Reinforcement Learning with Discrete World Models without Reconstruction

The present paper proposes a novel reinforcement learning method with wo...

Please sign up or login with your details

Forgot password? Click here to reset