On Bonus-Based Exploration Methods in the Arcade Learning Environment

by   Adrien Ali Taïga, et al.

Research on exploration in reinforcement learning, as applied to Atari 2600 game-playing, has emphasized tackling difficult exploration problems such as Montezuma's Revenge (Bellemare et al., 2016). Recently, bonus-based exploration methods, which explore by augmenting the environment reward, have reached above-human average performance on such domains. In this paper we reassess popular bonus-based exploration methods within a common evaluation framework. We combine Rainbow (Hessel et al., 2018) with different exploration bonuses and evaluate its performance on Montezuma's Revenge, Bellemare et al.'s set of hard of exploration games with sparse rewards, and the whole Atari 2600 suite. We find that while exploration bonuses lead to higher score on Montezuma's Revenge they do not provide meaningful gains over the simpler ϵ-greedy scheme. In fact, we find that methods that perform best on that game often underperform ϵ-greedy on easy exploration Atari 2600 games. We find that our conclusions remain valid even when hyperparameters are tuned for these easy-exploration games. Finally, we find that none of the methods surveyed benefit from additional training samples (1 billion frames, versus Rainbow's 200 million) on Bellemare et al.'s hard exploration games. Our results suggest that recent gains in Montezuma's Revenge may be better attributed to architecture change, rather than better exploration schemes; and that the real pace of progress in exploration research for Atari 2600 games may have been obfuscated by good results on a single domain.


page 16

page 18


Benchmarking Bonus-Based Exploration Methods on the Arcade Learning Environment

This paper provides an empirical evaluation of recently developed explor...

Improving Intrinsic Exploration with Language Abstractions

Reinforcement learning (RL) agents are particularly hard to train when r...

Are the results of the groundwater model robust?

De Graaf et al. (2019) suggest that groundwater pumping will bring 42–79...

Multi-Stage Episodic Control for Strategic Exploration in Text Games

Text adventure games present unique challenges to reinforcement learning...

Exploration by Random Network Distillation

We introduce an exploration bonus for deep reinforcement learning method...

Bayesian estimation of in-game home team win probability for National Basketball Association games

Maddox, et al. (2022) establish a new win probability estimation for col...

An exploration of the influence of path choice in game-theoretic attribution algorithms

We compare machine learning explainability methods based on the theory o...

Please sign up or login with your details

Forgot password? Click here to reset