Hyper-Parameter Sweep on AlphaZero General

03/19/2019
by   Hui Wang, et al.
6

Since AlphaGo and AlphaGo Zero have achieved breakground successes in the game of Go, the programs have been generalized to solve other tasks. Subsequently, AlphaZero was developed to play Go, Chess and Shogi. In the literature, the algorithms are explained well. However, AlphaZero contains many parameters, and for neither AlphaGo, AlphaGo Zero nor AlphaZero, there is sufficient discussion about how to set parameter values in these algorithms. Therefore, in this paper, we choose 12 parameters in AlphaZero and evaluate how these parameters contribute to training. We focus on three objectives (training loss, time cost and playing strength). For each parameter, we train 3 models using 3 different values (minimum value, default value, maximum value). We use the game of play 6×6 Othello, on the AlphaZeroGeneral open source re-implementation of AlphaZero. Overall, experimental results show that different values can lead to different training results, proving the importance of such a parameter sweep. We categorize these 12 parameters into time-sensitive parameters and time-friendly parameters. Moreover, through multi-objective analysis, this paper provides an insightful basis for further hyper-parameter optimization.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/12/2020

Analysis of Hyper-Parameters for Small Games: Iterations or Epochs in Self-Play?

The landmark achievements of AlphaGo Zero have created great research in...
research
02/15/2022

A Light-Weight Multi-Objective Asynchronous Hyper-Parameter Optimizer

We describe a light-weight yet performant system for hyper-parameter opt...
research
08/26/2022

Multi-objective Hyper-parameter Optimization of Behavioral Song Embeddings

Song embeddings are a key component of most music recommendation engines...
research
02/10/2021

Self-supervised learning for fast and scalable time series hyper-parameter tuning

Hyper-parameters of time series models play an important role in time se...
research
05/30/2017

Multi-Labelled Value Networks for Computer Go

This paper proposes a new approach to a novel value network architecture...
research
09/02/2010

Optimizing Selective Search in Chess

In this paper we introduce a novel method for automatically tuning the s...
research
05/25/2021

Reciprocal first-order second-moment method

This paper shows a simple parameter substitution, which makes use of the...

Please sign up or login with your details

Forgot password? Click here to reset